Search and filter through extracted claims from AI researchers.
Showing 41-60 of 278 claims in topic "benchmarks"
"we find that passage length is the main driver of performance, explaining 52-66% of the variance."
"the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task difficulty."
Google DeepMind is conducting the world's first double-blind AI evaluations
"Piloting the world's first double-blind AI evaluations"
"the proposed method yields a relative improvement of 23.6% on across-speaker ABX scores on average compared to non adapted models"
"We introduce ABX- Accent, a benchmark based on the AESRC dataset that features 10 different accents of English. It includes a small (< 10 hours) unlabelled training set in each of the accents and adaptations of the Zero Resources Challenge ABX evaluation metrics to each of the accents"
Compositional generalization is typically assessed using model accuracy as the primary metric
"Compositional generalization is usually evaluated through model accuracy."
"We instead ask which structural or lexical identifications make held-out COGS examples admissible from the structures observed in training."
"Across 21 COGS generalization types, admissibility follows distinct identification profiles, while residual failures separate unsupported structural templates."
"These data-side diagnoses characterize what the training corpus licenses under specified identifications, without training a predictive model."
"Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased."
"we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased."
PPSR yields more discriminative and stable metric rankings than existing system-level meta-metrics
"PPSR directly targets metric utility for prediction-powered evaluation and yields more discriminative and stable metric rankings than existing system-level meta-metrics."
"our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment"
The prediction-powered evaluation framework applies broadly to non-verifiable tasks
"applies broadly to non-verifiable tasks"
"We introduce AraMS-28k, the largest publicly released line-level dataset of genuine historical Arabic manuscripts, comprising 14 books, 3,043 pages, and 28,600 annotated text lines"
"margin lines that have an unambiguous attachment point in the main text are further annotated with an insertion anchor, recovering the manuscript's true non-linear reading order at line-level granularity -- to our knowledge the first such annotation released for a historical Arabic manuscript corpus"
"The dataset was constructed with RefLAM, a reference-grounded annotation pipeline that aligns multimodal-LLM OCR against independently sourced clean transcriptions and routes every line through human review, combining automatic verification with expert oversight"
"AraMS-28k is released with page images, line-level annotations, and fixed train/val/test splits under CC BY-NC-SA 4.0 to support reproducible research on Arabic manuscript recognition, layout analysis, and reading-order recovery"
"FID's first-two-moment summary can miss distributional differences"
"on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID $24.7$ versus $58.6$ for held-out real images (lower is better)"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.