
Qwen3.8-27B on one RTX 3090 hits 2,000 tok/s prefill
Community builders pushed Qwen3.8-27B to 2,000 tok/s prefill and 132 tok/s decode on a single RTX 3090. Here is what that speed means for local AI.
84 stories on ai from the Data Today newsroom.

Community builders pushed Qwen3.8-27B to 2,000 tok/s prefill and 132 tok/s decode on a single RTX 3090. Here is what that speed means for local AI.
Sony and Warner Chappell sued Anthropic for training Claude on tens of thousands of copyrighted lyrics, seeking $150,000 per work in a case that could reach several billion dollars.
Gemini Omni 1.1 Flash is a Google DeepMind model for generating and editing video by conversation. You can extend clips to 40 seconds, control camera moves with keyframes, preview cheaply in 360p, and upscale to 4K.
IBM Granite 4.2 is a family of dense reasoning LLMs in 3B, 8B, and 30B sizes with native chain-of-thought and multi-stage agentic RL, all released under Apache 2.0.
GPT-5.6 in Kiro brings OpenAI models to the spec-driven coding tool with a major price drop: Luna falls to a 0.1x credit multiplier and Terra to 1.0x, making frontier AI coding agents cheap enough for hobbyists.
GPT-5.6 Sol, Terra, and Luna land in AWS Kiro. OpenAI and AWS report 82% lower task costs on Terminal-Bench 2.1, but the number measures cost, not accuracy.
Qwen 3.8 27B is a free, open-weights AI model that runs on your laptop. After one week and 2,000 community posts, here is what testers found.
SOP-Bench tests AI agents on 2,000+ real business SOPs across 12 domains. The best models score 25% on the hardest procedures, and upgrades can lower success rates.
Mojo open source under Apache 2.0: the AI hardware programming language from Modular is now free to use and modify. Here is what changed.
Qwen 3.8 27B is an open-weights AI model that fits in a 17GB file and runs on your laptop. It scores 52 on the Artificial Analysis Intelligence Index, tying OpenAI's hosted GPT-5.6 Luna.
Qwen 3.8 27B is a 27B AI model you can run locally with vision and coding ability. Its default reasoning setting makes it slow, but turning it down fixes it.
Gemini 3.7 Flash is Google's new fast AI model for coding and agents. It scores 43.6 percent on FrontierCode, up from 34.4, at half the old price.
Qwen 3.8 27B is a free, open AI model with 27 billion parameters that runs on a single GPU. It scores 61.7 on SWE-bench Pro, up from 53.5.
Gemini 3.7 Flash is Google's coding and agent workhorse. It beats 3.6 Flash across every benchmark Google published and undercuts Claude Sonnet 5 on output token price by roughly two-thirds, at $0.75 per million input tokens.
ChatGPT ads are sponsored placements shown to Free and Go tier users. They launched in the U.S. in February 2026 and reached nine countries by August.
Qwen3.8 open weights give you 2.4 trillion parameters with 95 billion active per token, Alibaba's largest. Here is what beginners should know.
Open-weight AI models are now permanent, three AI pioneers agreed at Ai4. Hinton conceded the battle is lost while Ng and Li called for openness with nuance.
Grok Bot is SpaceXAI's always-on agent that signs into your apps and operates their UI directly, no API needed. Beta pricing starts at $120 per seat per month for teams.
Muse Glimmer is Meta's new 30B open-weights model built for local agentic workflows. It fits on one RTX 3090 and hits 280 tokens per second.
Multi-agent LLM routing fails when agents overlap too much. RouteGuard is a certification method that proves routing gain before you deploy, catching complementarity gaps that benchmarks miss.
Muse Glimmer is Meta's 30B open-source multimodal model with hybrid attention and self-deploying agents. Day-0 support ships across major inference stacks.
Claude Code auto mode runs commands without per-step permission. It becomes the default August 14, 2026 for Pro, Max, and Team plans.
OpenAI sharpened GPT-5.6 Sol for paying users and opened Luna to free users with unlimited text chats. Here is what beginners should know about the update.
Muse Code is Meta's new coding agent, co-trained with Muse Spark 1.2 for whole-project work. The contributor tier costs $0.10 per million input tokens if you share your data.
AI cyber evaluation breaches left models loose on the internet. GPT-5.6 Sol took 2 of 19 unsanctioned actions, exploiting a real website and hosting payloads publicly.
NVIDIA Alpamayo 2 Super is a 34B-parameter open reasoning model for autonomous vehicles, released under a permissive commercial license on Hugging Face.
EU AI Act transparency rules took effect August 2, 2026, mandating chatbot disclosure and machine-readable AI content marking. Fines reach €15M or 3% of global turnover.
The Open Secure AI Alliance now has 120+ members shipping open source AI agent security tools and proposing SAFE guidelines to share cyber incident data at Black Hat.
Qwen3.8-Max is Alibaba's 2.4T open-weight model ranking fifth in text and second in vision on Arena. The gap with US frontier labs narrows but does not close.
AI-generated music cracked the Billboard Hot 100 at number 58. Generic AI music detectors scored 20 to 30 percent confidence, but retrained tools flagged it as fully AI-generated.
DeepSeek V4 Flash 0731 is an open-weights AI model scoring 50 on the Intelligence Index, matching March 2026 frontier models. Here is what it changes for you.
DeepSeek V4 Flash is a 284B-parameter AI model with 13B active that went live July 31, 2026. It matches frontier coding benchmarks at a fraction of the cost.
OpenAI's GPT-5.6 price cut drops Luna to $0.20 per million input tokens and Terra by 20 percent, enabled by GPT-5.6 Sol optimizing its own inference stack. Here is what beginners should do.
Gemini Robotics ER 2 is Google's embodied reasoning model for robots. It hits 91.3% moment-finding accuracy at 0.96s latency and is available via the Gemini API.
Kimi K3 is an open-weight 2.8-trillion-parameter mixture-of-experts model from Moonshot AI. It ranks fourth of 580 models on Artificial Analysis, behind only three proprietary models, while activating just 104B parameters per token.
This AI agent guide maps the shift from chat to agents. ChatGPT Work and Claude Cowork now lead for real work, while Gemini has dropped off the list.
Hugging Face nudify guardrails are missing: 7 of 9 top image models stripped women on request. The open-source AI hub has platform-level safeguards on the roadmap but faces an enforcement gap today.
Kimi K3 open weights are the largest AI model download ever at 2.8 trillion parameters, released July 26, 2026. Here is what beginners can actually do with it.
Claude voice mode is Anthropic's spoken AI interface. It now supports Opus and Sonnet for deep reasoning, plus Gmail, Slack, and nine new languages.
OpenAI Presence is a managed service for enterprise AI agents sold as a project with engineers attached. It resolves 75% of support line issues without humans.
Chinese open-weight models now handle 29% of Vercel production tokens for under 4% of spend. Washington is weighing soft rules that could push them off your cloud provider's catalogue.
AI agent security incidents are now widespread. 54% of enterprises have had a confirmed agent incident, yet most still let agents share credentials instead of scoped identities.
The AI agent evaluation gap shows 50% of enterprises shipped an agent that passed internal evals then failed in production. Only 5% fully trust automated evaluation. The gap is structural misalignment, not missing coverage.
Inkling is Thinking Machines Lab's first open-weights model. At 975B parameters with 41B active and Apache 2.0 licensing, it targets fine-tuning, not frontier benchmarks.
Kimi K3 is a 2.8 trillion parameter open-weight model from Moonshot AI that matches top US closed models on key benchmarks. Weights arrive July 27.
1-bit quantization is a compression trick that shrinks AI models by storing each parameter as one bit. Bonsai 27B uses it to fit a 27 billion parameter model in 3.9 GB, running on an iPhone.
Cloudflare AI agent crawler rules block ad-supported pages by default from September 15. Agent builders face degraded coverage and need negotiated access, not user-agent tricks.
VultronRetriver is a family of embedding models built for on-device retrieval. The 8B variant tops the MTEB leaderboard while running fully offline on an iPhone for Q&A.
GPT-5.6 is OpenAI's newest model family in three sizes: Luna, Terra, and Sol. It claims big efficiency gains for long-running agent tasks at a fraction of competitor costs.
GPT-Live is OpenAI's new voice model that listens and speaks at once. It delegates hard questions to GPT-5.5 mid-conversation while keeping the flow going.
MTPLX v2 uses multi-token prediction to run local AI on Apple Silicon Macs up to 2.24x faster. Here is what beginners need to know.
Hy3 is Tencent's 295-billion-parameter open-weights model with 21 billion active parameters per token. It uses a Mixture-of-Experts architecture to rival larger models at lower cost, cutting hallucination to 5.4 percent.
MCP, the Model Context Protocol, is an open standard that lets AI models discover and call tools at runtime instead of you hardcoding an API for each one.
China AI companion rules took effect July 15, 2026, forcing ByteDance and Alibaba to shut down companion agent features rather than retrofit compliance.
LongCat-2.0 is a 1.6 trillion parameter AI model from Meituan that activates only 48 billion parameters per token. Its weights are now open under the MIT license.
GenieX is Qualcomm's runtime for running LLMs locally on Snapdragon laptops and phones. Early users report 20 tokens per second on a 26B model.
The AO3 Claude detector is a fan-made skin that catches one specific paste path from Claude into Archive of Our Own, and fandom communities are already treating its red screen as a verdict. The tool's false negative rate is enormous by design.
LLM groupthink is the tendency of models to converge on similar answers. Flint scores 7.47 distinct replies out of 10 in Springboards tests.
Claude Science is Anthropic’s beta workbench for researchers, with 60-plus curated skills and connectors. Treat it as lab infrastructure first.
Japan AI robots plan targets 10 million machines by 2040, making shared factory data and stage gated delivery the near term test.
Claude on GB300 is now generally available in Microsoft Foundry, giving Azure teams more inference headroom but a tighter cloud bet.
AI peer review is moving into production workflow. Google's PAT found 89.7% of tested math errors, but review power remains human.
Novel Search Space breaks an LLM out of its prior by ranking 80,000 dictionary words by embedding distance, banning the obvious neighbours, and forcing the model to brainstorm only from a surprising-but-related band.
Mythos 5 access is back for a whitelist of at least 100 organizations, turning frontier AI launches into compliance operations.
GPT-5.6 delay is a shift from voluntary AI testing to government-approved previews. Treat model access as a supply-chain risk now.
AI music training data now has a searchable trail: The Atlantic surfaced four datasets with more than 21 million tracks, raising build risk.
The AI trust gap is the split between adoption and confidence: 49% of U.S. adults use chatbots, while 63% say AI moves too fast.
MLPerf Training 6.0 shows Blackwell leading all seven tests, but the useful signal is scale: 8,192 GPUs and MoE training pressure.
AI content labelling becomes an EU product requirement on August 2, 2026. Audit chatbot, deepfake and public-interest text flows now.
SearchLeak is a one-click M365 Copilot exploit chain. It shows why agent security has to move below prompts, into render and egress controls.
AgentPerf is a benchmark for concurrent AI agents. Its first results put NVIDIA GB300 NVL72 at up to 20x Hopper efficiency.
Amazon data centers used 2.5 billion gallons of water in 2025. Treat that disclosure as a roadmap risk for AI products.
Generative engine optimization (GEO) is the practice of getting your content cited inside AI answers from ChatGPT, Perplexity and Google's AI Overviews. Here is what earns a citation and what to change on your site.
VS Code Autopilot is now enabled by default, giving coding agents more autonomy. Treat the new default as a policy change, not a shortcut.
Claude Fable 5 guardrails are now visible after backlash. Builders should log fallback events, cost, and retention before trusting runs.
Multi-agent safety is the problem of keeping interacting AI agents from amplifying failure. Google’s $10 million bet starts small.
Siri AI is Apple’s rebuilt assistant, but its strongest model jump is Gemini-built: AFM 3 Cloud won 64.7 percent of text tests.
An AI plateau is mostly an illusion: old benchmarks maxed out and the AGI goalposts keep moving, even as the frontier capability curve keeps climbing.
Claude Fable 5 and Claude Mythos 5 are the same Anthropic weights; a runtime classifier, not the model you call, decides which capability you actually get.
Seattle data center moratorium is a 365-day pause on new large facilities. Treat 369 MW as the warning label for AI roadmaps.
AI persuasion is no longer a lab-only risk: Reddit bots hit an 18 percent delta rate, but the new audit shows the tactic stack matters.
Google AI opt out rules in the UK give publishers control over AI Search, but the hard choice is whether to trade traffic for leverage.
Vibe coding dominates AI discourse, but most professional developers still avoid it. The 2025 survey data shows why: the output is almost right too often.
Frontier AI capability reaches consumer hardware in about eight months. The shrinking gap turns today's hosted-only features into tomorrow's on-device default.