Search and filter through extracted claims from AI researchers.
Showing 61-80 of 278 claims in topic "benchmarks"
"FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged and therefore do not encode the direction of a dispersion change: under-dispersion, as can occur in mode collapse, versus over-dispersion"
"ZID reports three linked outputs: an index for ranking departure magnitude, a permutation $p$-value for testing distributional equality, and a signed dispersion readout for diagnosis"
"In controlled experiments, ZID detects a broad range of departures, and its score tracks increasing severity along the corresponding sweeps, including cases in which FID is flat or reversed"
"On DiT-XL/2 and SiT-XL/2 guidance sweeps, ZID detects departure from real data, and its signed readout labels the high-guidance diversity collapse as under-dispersion"
"The rate of vulnerabilities reported across many projects has dramatically accelerated in 2026 compared with 2025, both for specific projects (cURL, OpenSSL, Firefox, and Microsoft) and for aggregate vulnerability databases (the US NVD, and OSV)"
"AI is clearly contributing to more work being done (arXiv submissions have doubled in some areas in less than 12 months) but quantifying the value of those contributions is difficult"
"When you look at algorithmic progress across seven significant problem areas (CIFAR-10, Hutter compression, Gurobi mixed-integer programming, MIPLIB, nanoGPT, Stockfish, and the matrix-multiplication exponent) there are a couple of these where LLM-attributable contributions have happened (nanoGPT, CIFAR-10), though the rate of increase of usage of AI here is a lot less than with cybersecurity and mathematics"
"My suspicion is that acceleration happens when models go through some kind of ineffable phase change for a given skill, as has evidently happened with day-to-day coding (2025), and cyber (2026)"
"at 30B-A3B, SPADE reaches a suite average of 58.3: +8.1 over base and +5.3 over the strongest fixed-environment baseline"
Opus 5 and Fable 5 with Claude Code are the best overall models on DiG-bench
"Opus 5 and Fable 5 with Claude Code are the best overall models, followed by GPT-5.5"
Only Opus 5 and Fable 5 were able to beat any tasks in Tier 7 of DiG-bench
"Only Opus 5 and Fable 5 were able to beat any tasks (0.2) in (Tier 7)."
"This puts the model more or less at the frontier of agentic coding benchmarks, with only ~750B parameters – a third of Kimi K3!"
"AI checkers are essentially a cat-and-mouse game. AI checkers may learn to detect a certain pattern that is indicative of AI-generated content. Then, the next LLM may incidentally or deliberately not exhibit that pattern and avoid detection. The AI checker then has to be updated to detect said LLM, and so forth."
"However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation provides limited diagnostic value for identifying process-level failures or rigorous logic, failing to guide the transformation of LLMs into robust agents."
Models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles
"Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles."
"This demonstrates that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.