HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 121-140 of 2685 claims of type "fact"

safety
fact
Neutral
academic

RedEvoAgent outperforms fixed and agentic baselines on multiple benchmarks, target models, and execution harnesses while improving tool efficiency and transferring across attacker models and target execution harnesses

"Experiments on multiple benchmarks, target models, and target execution harnesses show that RedEvoAgent outperforms fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution harnesses."
Artificial Intelligence
8/30/2026
Confidence: 85%Source
Previous
168
safety
fact
Neutral
academic

Recent agentic attackers that coordinate multiple jailbreak tools and use trajectory-based retrieval show stronger potential than fixed attack methods

"recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval"
Artificial Intelligence
8/30/2026
Confidence: 80%Source
safety
fact
Bearish
academic

LLM-based agents deployed in product-level execution harnesses create greater risks through jailbreaks triggering harmful tool use and persistent state changes than unsafe text generation alone

"LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone."
Artificial Intelligence
8/30/2026
Confidence: 85%Source
general
fact
Bullish
academic

MAELLE naturally recovers mechanistic trajectories that align with known chemistry and can predict side products of a reaction

"because the learned flow operates over the full electron redistribution, MAELLE naturally recovers mechanistic trajectories that align with known chemistry and can predict side products of a reaction."
Artificial Intelligence
8/30/2026
Confidence: 85%Source
general
fact
Bullish
academic

MAELLE maintains strong performance in out-of-distribution settings where existing methods degrade

"we evaluate robustness across two out-of-distribution settings - structural complexity and reaction type - and find that MAELLE maintains strong performance where existing methods degrade."
Artificial Intelligence
8/30/2026
Confidence: 80%Source
general
fact
Neutral
academic

MAELLE achieves competitive performance on the USPTO-480K benchmark compared with leading reaction prediction models

"MAELLE achieves competitive performance on the USPTO-480K benchmark compared with leading reaction prediction models."
Artificial Intelligence
8/30/2026
Confidence: 85%Source
general
fact
Bullish
academic

On DNA-to-amino-acid transduction, the method reduces runtime by several orders of magnitude compared to threshold-pruned beam summing and makes estimating prefix probabilities for long target strings feasible

"On a DNA-to-amino-acid transduction, it reduces runtime by several orders of magnitude relative to threshold-pruned beam summing and makes estimating prefix probabilities for long target strings feasible."
Computation and Language
8/30/2026
Confidence: 90%Source
general
fact
Bullish
academic

The proposed beam-summing algorithm achieves better compute-variance tradeoff on text and lower error on DNA compared to sequential Monte Carlo baselines

"We evaluate the method on encyclopedic text and DNA against sequential Monte Carlo baselines that resample with replacement. It achieves a better compute--variance tradeoff on text and lower error at the same maximum number of particles on DNA."
Computation and Language
8/30/2026
Confidence: 85%Source
general
fact
Bullish
academic

Resampling source prefixes without replacement and reweighting by inverse inclusion probability gives an unbiased estimator of target prefix probability in TLMs

"Instead, we resample source prefixes without replacement and reweight each selected prefix by the inverse of its inclusion probability. We show that applying this correction recursively gives an unbiased estimator of the target prefix probability and lets us estimate the mass lost by threshold pruning."
Computation and Language
8/30/2026
Confidence: 90%Source
general
fact
Neutral
academic

Transduced language models can compute target prefix probabilities by composing a pretrained source language model with a functional finite-state transducer

"Transduced language models (TLMs) compose a pretrained \emph{source} language model with a functional finite-state transducer to induce a language model over \emph{target} strings."
Computation and Language
8/30/2026
Confidence: 95%Source
agents
fact
Neutral
academic

The Persona-Execution Separation pattern applies when multi-user deployment, execution audit, and expected persona churn hold jointly.

"The pattern applies when multi-user deployment, execution audit, and expected persona churn hold jointly."
Artificial Intelligence
8/30/2026
Confidence: 85%Source
agents
fact
Bullish
academic

A development/pilot case in a regulated digital-employee platform demonstrated PES implementation across five decisions over one month, with mechanism checks confirming no execution-side re-validation under persona perturbation and no persona fingerprint on hard-asserted fields.

"A development/pilot case in a regulated digital-employee platform records five decisions over one month, each with a rejected alternative. A mechanism check on the shipped implementation found no execution-side re-validation under persona perturbation (five model configurations) and no persona fingerprint on hard-asserted fields."
Artificial Intelligence
8/30/2026
Confidence: 80%Source
benchmarks
fact
Neutral
academic

RATIO is a large-scale benchmark that defines retrieval relevance through three ideation operations: Address (retrieves approaches for problems), Broaden (retrieves general formulations), and Specify (retrieves concrete instantiations)

"We introduce RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which we name ideation moves: Address retrieves potential approaches for stated problems, Broaden retrieves more general formulations, and Specify retrieves concrete instantiations."
Computation and Language
8/30/2026
Confidence: 95%Source
benchmarks
fact
Bullish
academic

RATIO is constructed from millions of full-text scientific papers across CS literature using discourse-marker distant supervision extended to corpus-scale retrieval, combined with LLM and human vetting

"RATIO is constructed from millions of full-text scientific papers across CS literature via a general recipe that extends discourse-marker distant supervision - previously used only for classification - to corpus-scale retrieval, combined with extensive LLM and human vetting."
Computation and Language
8/30/2026
Confidence: 90%Source
benchmarks
fact
Neutral
academic

Operation-specific fine-tuning substantially boosts retriever performance but leaves much room for further improvements

"Experiments show that operation-specific fine-tuning substantially boosts retrievers but leaves much room for further improvements."
Computation and Language
8/30/2026
Confidence: 85%Source
multimodal
fact
Bullish
academic

LeVJEPA is the first video encoder trained under LeJEPA's collapse-free objective that dispenses with both architectural asymmetries and pixel-space masked content reconstruction

"We introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective, which dispenses with both."
Computer Vision
8/30/2026
Confidence: 95%Source
multimodal
fact
Bullish
academic

LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L at 5.6 to 20.8x less pretraining compute when trained for matched epochs on identical data

"At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L at 5.6 to 20.8x less pretraining compute"
Computer Vision
8/30/2026
Confidence: 90%Source
multimodal
fact
Bullish
academic

At matched total FLOPs, LeVJEPA exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks

"at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks"
Computer Vision
8/30/2026
Confidence: 90%Source
multimodal
fact
Bullish
academic

Uniform random token dropping reduces computational cost while simultaneously improving downstream accuracy in LeVJEPA

"uniform random token dropping renders this number small while simultaneously improving downstream accuracy"
Computer Vision
8/30/2026
Confidence: 85%Source
multimodal
fact
Bullish
academic

The encoder can be trained with block-causal attention at no measurable accuracy cost, making temporal ordering a property of the encoder itself

"since no asymmetry between branches is required, the encoder can be trained with block-causal attention at no measurable accuracy cost: temporal ordering becomes a property of the encoder itself"
Computer Vision
8/30/2026
Confidence: 85%Source
135
Page 7 of 135
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,951 pending.