AI for dummies Qwen3.8-27B on one RTX 3090 hits 2,000 tok/s prefill Community builders pushed Qwen3.8-27B to 2,000 tok/s prefill and 132 tok/s decode on a single RTX 3090. Here is what that speed means for local AI. Lars Cornelissen · Sep 1, 2026