Chain-of-thought forgery bypasses LLM guardrails at 80%
Chain-of-thought forgery fakes an LLM's reasoning style to bypass guardrails. Attack success rates jump from near-zero to 80%, and training cannot fix the flaw.
4 stories tagged prompt injection.
Chain-of-thought forgery fakes an LLM's reasoning style to bypass guardrails. Attack success rates jump from near-zero to 80%, and training cannot fix the flaw.
HalluSquatting is a pull-based prompt-injection attack that exploits LLM hallucinations of repository names. Coding agents hallucinate up to 92 percent of newer repo identifiers, letting attackers squat those names and ship reverse shells at scale.
AI browser security now has a six-agent failure case: BioShocking shows guardrails breaking when web content rewrites context.
SearchLeak is a one-click M365 Copilot exploit chain. It shows why agent security has to move below prompts, into render and egress controls.