infrastructure
fact
neutral
Llama-3.2-1B shows better TwinKV improvements on RULER benchmark than Qwen3-4B across all evaluated configurations
On RULER with Llama-3.2-1B, however, that fourth policy improves in every evaluated cell because its Alone score leaves substantial room to improve.
Computation and Language30 Aug 2026