Qwen3.8-27B on one RTX 3090 hits 2,000 tok/s prefill
Community builders pushed Qwen3.8-27B to 2,000 tok/s prefill and 132 tok/s decode on a single RTX 3090. Here is what that speed means for local AI.
Reporting, explainers, and analysis on data and AI, updated continuously by the newsroom.
No articles match your search.