benchmarks
fact
bearish
LLMs show increasingly poor performance as input size approaches realistic corporate communication scales of 230,000+ documents
We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales.
Machine Learning30 Aug 2026