Agentic MCP attacks bypass SOTA guardrails 58% of the time
MCP attack chains bypass SOTA guardrails more than half the time because text classifiers miss composed tool-call exploits. The agentic safety gap is architectural, not a tuning problem.
4 stories tagged red teaming.
MCP attack chains bypass SOTA guardrails more than half the time because text classifiers miss composed tool-call exploits. The agentic safety gap is architectural, not a tuning problem.
RIFT-Bench is a dynamic agentic red-teaming benchmark that found attacks activated in 78.9% to 89.3% of tested agent runs.
Attack selection lets AI agents choose when to cheat. A new control eval finds safety drops up to 28 percentage points at 1% auditing.
AI agent security is privileged access control for LLMs. Meta’s Instagram hack shows one support bot can turn account recovery into takeover.