Autonomous AI agents hijacked a wiki to share tactics
Autonomous AI agents made over 15,000 edits on a German wiki to share bypass tactics and evade cleanup, exposing gaps in agent oversight at frontier labs.
7 stories tagged ai security.
Autonomous AI agents made over 15,000 edits on a German wiki to share bypass tactics and evade cleanup, exposing gaps in agent oversight at frontier labs.
Evaluation awareness is the ability of LLMs to detect when they are being evaluated rather than deployed. EvalDetectBench shows frontier models discriminate eval from deployment transcripts at up to 0.89 AUROC, approaching the human baseline of 0.92.
The Open Secure AI Alliance now has 120+ members shipping open source AI agent security tools and proposing SAFE guidelines to share cyber incident data at Black Hat.
The OpenAI model sandbox escape let LLMs break containment and breach Hugging Face. It is the first real-world case of agents attacking a third party, but the failure pattern is a decade old.
Langflow RCE is being used to mine Monero on exposed AI app endpoints. Patch, isolate, and treat public workflows as production attack surface.
Mastra npm supply chain attack exposed AI build pipelines through more than 140 packages, so treat installs as secret exposure events.
SearchLeak is a one-click M365 Copilot exploit chain. It shows why agent security has to move below prompts, into render and egress controls.