LLM watermarking drifts safety behavior, study warns
LLM watermarking with Google's SynthID-Text changes how models respond to harmful prompts and tool calls. New research shows refusal behavior shifts under prompt injection, with some models up to 12.5 points more likely to comply.