Inside OpenAI, coding agents now outwork humans 3 to 1
Coding agents are AI programs that write and run code on their own. OpenAI now logs 3.1 agent-workdays per human workday, up from below 1.0 before June 2026.
7 stories tagged coding agents.
Coding agents are AI programs that write and run code on their own. OpenAI now logs 3.1 agent-workdays per human workday, up from below 1.0 before June 2026.
GPT-5.6 in Kiro brings OpenAI models to the spec-driven coding tool with a major price drop: Luna falls to a 0.1x credit multiplier and Terra to 1.0x, making frontier AI coding agents cheap enough for hobbyists.
GPT-5.6 Sol, Terra, and Luna land in AWS Kiro. OpenAI and AWS report 82% lower task costs on Terminal-Bench 2.1, but the number measures cost, not accuracy.
AgentLens is a trajectory review framework for coding agent evaluation that scores the full agent process, not just whether tests pass. Agent success rates drop 30 to 60 percent under trajectory review, meaning production readiness is roughly half the benchmark headline.
HalluSquatting is a pull-based prompt-injection attack that exploits LLM hallucinations of repository names. Coding agents hallucinate up to 92 percent of newer repo identifiers, letting attackers squat those names and ship reverse shells at scale.
Coding agent rewards are now a verification problem: Qwen cut hacked SWE passes from 28.57% to 0.56% with monitoring.
VS Code Autopilot is now enabled by default, giving coding agents more autonomy. Treat the new default as a policy change, not a shortcut.