
Qwen3.8-27B on one RTX 3090 hits 2,000 tok/s prefill
Community builders pushed Qwen3.8-27B to 2,000 tok/s prefill and 132 tok/s decode on a single RTX 3090. Here is what that speed means for local AI.
Beginner-friendly apps, plugins and coding copilots worth your first hour, and what they actually cost to run.

Community builders pushed Qwen3.8-27B to 2,000 tok/s prefill and 132 tok/s decode on a single RTX 3090. Here is what that speed means for local AI.
Mojo open source under Apache 2.0: the AI hardware programming language from Modular is now free to use and modify. Here is what changed.
Claude Code auto mode runs commands without per-step permission. It becomes the default August 14, 2026 for Pro, Max, and Team plans.
This AI agent guide maps the shift from chat to agents. ChatGPT Work and Claude Cowork now lead for real work, while Gemini has dropped off the list.
MTPLX v2 uses multi-token prediction to run local AI on Apple Silicon Macs up to 2.24x faster. Here is what beginners need to know.
GenieX is Qualcomm's runtime for running LLMs locally on Snapdragon laptops and phones. Early users report 20 tokens per second on a 26B model.