Search and filter through extracted claims from AI researchers.
Showing 1-20 of 278 claims in topic "benchmarks"
"We introduce RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which we name ideation moves: Address retrieves potential approaches for stated problems, Broaden retrieves more general formulations, and Specify retrieves concrete instantiations."
"RATIO is constructed from millions of full-text scientific papers across CS literature via a general recipe that extends discourse-marker distant supervision - previously used only for classification - to corpus-scale retrieval, combined with extensive LLM and human vetting."
"Experiments show that operation-specific fine-tuning substantially boosts retrievers but leaves much room for further improvements."
"RATIO provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scientific inspiration retrieval."
"We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales."
Current synthetic datasets for evaluating LLM document understanding have been overly simple
"synthetic datasets have been overly simple"
"LLMs are increasingly able to answer complex questions about enterprise-scale document collections."
"CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem."
"it is unclear whether existing AI systems are inclusive enough for blind and deafblind users to access the same functionality through Braille"
There is a persistent gap between LLM capabilities in print-English versus Braille accessibility
"The results reveal a persistent gap between print-English capability and Braille accessibility"
"Braille understanding and expression are asymmetric, where Grade 2 is especially fragile on the input side compared to Grade 1"
Fully Braille requests further reduce LLM performance beyond the existing accessibility gap
"fully Braille requests further reduce performance"
BrailleBench provides valuable guidance for the development of future Braille AI systems
"The experimental observations provide valuable guidance for the development of future Braille AI systems"
"We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation"
"a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them"
"The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$)"
"The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone"
"We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings"
"High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little."
"This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,947 pending.