GPT-6 Astra benchmarks are strong, but watch the caveats
GPT-6 Astra is OpenAI's latest flagship model. It scores 98.6% on ARC-AGI-3 and tops most benchmarks, but costs 2.5x Sol and splits evaluators on real gains.
5 stories tagged agentic ai.
GPT-6 Astra is OpenAI's latest flagship model. It scores 98.6% on ARC-AGI-3 and tops most benchmarks, but costs 2.5x Sol and splits evaluators on real gains.
Copilot billing shock is the sticker-price moment for agentic coding. Per-token costs hit $0.07 and enterprise invoices spike as agents loop.
MCP attack chains bypass SOTA guardrails more than half the time because text classifiers miss composed tool-call exploits. The agentic safety gap is architectural, not a tuning problem.
Agentic AI work is moving from chat to delegated tasks: OpenAI says 70.2% of sampled Codex users handed off one hour of work.
AgentPerf is a benchmark for concurrent AI agents. Its first results put NVIDIA GB300 NVL72 at up to 20x Hopper efficiency.