Research Declarative Attention cuts LLM decoding tokens by 52% Declarative Attention lets LLMs declare which context regions to read, cutting attended tokens 52% with only 1.27pp accuracy loss on Gemma-4-31B. Lars Cornelissen · Sep 5, 2026