Engineering
A two-bit LLM runs on a laptop in 60 MB and hits 400 tok/s
Sub-2-bit LLM quantization puts a 250M model in 60 MB of disk and 80 MB of RAM, running at 400 tok/s on a laptop CPU. The method keeps recent tokens at full precision and compresses the rest, pointing toward a new class of lightweight deployments.