benchmarks
critique
bearish
Existing Indonesian cultural commonsense benchmarks fail to capture cultural nuances because they evaluate LLMs on short, isolated prompts rather than dialogic contexts where culture is actually lived
Computation and Language28 Jul 2026