benchmarkscritiqueneutralCurrent synthetic datasets for evaluating LLM document understanding have been overly simplesynthetic datasets have been overly simpleMachine Learning30 Aug 2026http://arxiv.org/abs/2608.27391v1