Engineering
Ising glass LLM block pruning saves 23 MMLU points at 50%
LLM block pruning via Ising glass optimization preserves MMLU near 77 at 50% compression of Llama-3.3-70B-Instruct, beating baselines by 23 percentage points.
2 stories tagged llm-compression.
LLM block pruning via Ising glass optimization preserves MMLU near 77 at 50% compression of Llama-3.3-70B-Instruct, beating baselines by 23 percentage points.
Quantization-Aware Healing distills 4-bit compressed LLMs from the original teacher. The 4-bit student beats its bfloat16 source on 7 of 9 benchmarks.