Research
Declarative Attention cuts LLM decoding tokens by 52%
Declarative Attention lets LLMs declare which context regions to read, cutting attended tokens 52% with only 1.27pp accuracy loss on Gemma-4-31B.
1 story tagged declarative attention.