Search and filter through extracted claims from AI researchers.
Showing 101-120 of 278 claims in topic "benchmarks"
A single scalar success rate cannot adequately explain computer-use agent performance
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.