benchmarks
fact
neutral
Large language models are strong on knowledge-intensive topics like history, geography, and mathematics, but substantially weaker on everyday popular-culture topics such as celebrities, music, movies, and news
Computation and Language29 Jul 2026