benchmarks
critique
neutral
Existing benchmarks for long-term memory in LLMs remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning
Computation and Language29 Jul 2026