Search and filter through extracted claims from AI researchers.
Showing 201-220 of 278 claims in topic "benchmarks"
Muse Spark 1.1 beats all competitor models except Fable/Mythos on HealthBench-Pro
GPT-5.6 has been optimized specifically for benchmark performance (benchmaxxed)
Muse Spark 1.1 achieves +5% better performance than Muse Spark 1.0 on HealthBench-Pro
SWE-Bench Pro is now saturated/terminally flawed according to OpenAI's evals team
Two benchmarks suggest AI capabilities have been improving exponentially in recent months
Test-time compute budgets have significant impact on frontier AI model evaluations
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.