Search and filter through extracted claims from AI researchers.
Showing 1-20 of 178 claims in topic "benchmarks" of type "fact"
"We introduce RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which we name ideation moves: Address retrieves potential approaches for stated problems, Broaden retrieves more general formulations, and Specify retrieves concrete instantiations."
"RATIO is constructed from millions of full-text scientific papers across CS literature via a general recipe that extends discourse-marker distant supervision - previously used only for classification - to corpus-scale retrieval, combined with extensive LLM and human vetting."
"Experiments show that operation-specific fine-tuning substantially boosts retrievers but leaves much room for further improvements."
"We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales."
"LLMs are increasingly able to answer complex questions about enterprise-scale document collections."
"it is unclear whether existing AI systems are inclusive enough for blind and deafblind users to access the same functionality through Braille"
There is a persistent gap between LLM capabilities in print-English versus Braille accessibility
"The results reveal a persistent gap between print-English capability and Braille accessibility"
"Braille understanding and expression are asymmetric, where Grade 2 is especially fragile on the input side compared to Grade 1"
Fully Braille requests further reduce LLM performance beyond the existing accessibility gap
"fully Braille requests further reduce performance"
"We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation"
"a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them"
"The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$)"
"The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone"
"We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings"
"High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little."
R2M-Bench's Overall NMR metric correlates with human consistency judgments at Spearman's ρ=0.547
"Overall NMR correlates with human consistency judgments at Spearman's $ρ=0.547$ (95\% CI $[0.45,0.63]$)."
"Its within-model correlation magnitude with generated motion is $0.072$, compared with $0.207$ for raw revisit similarity, indicating that relative calibration substantially reduces the slow-motion shortcut."
DreamX-World-Memo achieves the highest Overall NMR among evaluated video models
"DreamX-World-Memo achieves the highest Overall NMR among the evaluated video models."
"Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks."
"The 532-row release adds clarification, goal revision, timestamp policy, quality-gated reporting, and evidence attribution while preserving the source computation and split."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.