KVarN: Variance-Normalized KV-Cache Quantization [R]

📰 Reddit r/MachineLearning

Excited to share some of my own work here :) KVarN is our new KV-Cache quantization method. In very brief, we combine Hadamard rotations with variance-normalization on both axes of the K and V matrices, then round to nearest. Simple, but works very well, especially for decode-heavy test-time-scaling settings (reasoning, code-gen, agentics). We get 3-4x compression at virtually no accuracy drop (mostly 0-1%) on tough benchmarks l

Published 4 Jun 2026
Read full article → ← Back to Reads