HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 1-20 of 278 claims in topic "benchmarks"

benchmarks
fact
Neutral
academic

RATIO is a large-scale benchmark that defines retrieval relevance through three ideation operations: Address (retrieves approaches for problems), Broaden (retrieves general formulations), and Specify (retrieves concrete instantiations)

"We introduce RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which we name ideation moves: Address retrieves potential approaches for stated problems, Broaden retrieves more general formulations, and Specify retrieves concrete instantiations."
Computation and Language
8/30/2026
Confidence: 95%Source
214
Page 1 of 14Next
benchmarks
fact
Bullish
academic

RATIO is constructed from millions of full-text scientific papers across CS literature using discourse-marker distant supervision extended to corpus-scale retrieval, combined with LLM and human vetting

"RATIO is constructed from millions of full-text scientific papers across CS literature via a general recipe that extends discourse-marker distant supervision - previously used only for classification - to corpus-scale retrieval, combined with extensive LLM and human vetting."
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Neutral
academic

Operation-specific fine-tuning substantially boosts retriever performance but leaves much room for further improvements

"Experiments show that operation-specific fine-tuning substantially boosts retrievers but leaves much room for further improvements."
Computation and Language
8/30/2026
Confidence: 85%Source
benchmarks
opinion
Bullish
academic

RATIO opens up new research avenues on scientific inspiration retrieval by providing a scalable training and evaluation framework for retrieval components that support literature-grounded ideation

"RATIO provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scientific inspiration retrieval."
Computation and Language
8/30/2026
Confidence: 80%Source
benchmarks
fact
Bearish
academic

LLMs show increasingly poor performance as input size approaches realistic corporate communication scales of 230,000+ documents

"We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales."
Machine Learning
8/30/2026
Confidence: 85%Source
benchmarks
critique
Neutral
academic

Current synthetic datasets for evaluating LLM document understanding have been overly simple

"synthetic datasets have been overly simple"
Machine Learning
8/30/2026
Confidence: 90%Source
benchmarks
fact
Bullish
academic

LLMs are increasingly capable of answering complex questions about enterprise-scale document collections

"LLMs are increasingly able to answer complex questions about enterprise-scale document collections."
Machine Learning
8/30/2026
Confidence: 75%Source
benchmarks
opinion
Neutral
academic

There is a crucial gap in the benchmarking ecosystem for corporate communication reasoning that needs to be filled

"CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem."
Machine Learning
8/30/2026
Confidence: 80%Source
benchmarks
fact
Bearish
academic

Existing AI systems have unclear inclusivity for blind and deafblind users accessing functionality through Braille

"it is unclear whether existing AI systems are inclusive enough for blind and deafblind users to access the same functionality through Braille"
Computation and Language
8/30/2026
Confidence: 80%Source
benchmarks
fact
Bearish
academic

There is a persistent gap between LLM capabilities in print-English versus Braille accessibility

"The results reveal a persistent gap between print-English capability and Braille accessibility"
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Bearish
academic

Braille understanding and expression are asymmetric in LLMs, with Grade 2 being especially fragile on the input side compared to Grade 1

"Braille understanding and expression are asymmetric, where Grade 2 is especially fragile on the input side compared to Grade 1"
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Bearish
academic

Fully Braille requests further reduce LLM performance beyond the existing accessibility gap

"fully Braille requests further reduce performance"
Computation and Language
8/30/2026
Confidence: 85%Source
benchmarks
opinion
Bullish
academic

BrailleBench provides valuable guidance for the development of future Braille AI systems

"The experimental observations provide valuable guidance for the development of future Braille AI systems"
Computation and Language
8/30/2026
Confidence: 70%Source
benchmarks
fact
Bearish
academic

Double-difference designs used to audit LLM judge bias are not properly identified on bounded rating scales, confounding differential preference with differential attenuation

"We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation"
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Bearish
academic

A severity shift common to both responses can manufacture a spurious interaction effect whenever two responses censor it unequally due to unequal distances from scale bounds

"a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them"
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Neutral
academic

In a pre-registered audit of a pedagogy judge with 990 calls, the effect of stated learner profile on scaffolding preference was null at +0.085 points

"The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$)"
Computation and Language
8/30/2026
Confidence: 95%Source
benchmarks
fact
Bearish
academic

A nominally significant interaction effect of +0.378 in the audit is not actually identified as preference, with 79-85% of it reproducible from severity shift and scale floor alone without any differential preference

"The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone"
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Neutral
academic

The contribution of the censoring mechanism to observed interaction effects is measurable from an audit's own ratings and can be derived in closed form

"We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings"
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Neutral
academic

High similarity between first-visit and return frames in video world models does not necessarily indicate scene memory, as the intervening rollout may have changed very little

"High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little."
Computer Vision
8/30/2026
Confidence: 85%Source
benchmarks
critique
Bearish
academic

Absolute revisit scores are sensitive to rendering stability, repetitive content, and failed motion, making them unreliable metrics

"This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion."
Computer Vision
8/30/2026
Confidence: 80%Source

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,947 pending.