HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 141-160 of 637 claims in topic "safety"

safety
fact
Bearish
academic

I2V systems have a temporal vulnerability where unsafe semantics can emerge from semantic composition over time rather than from a single frame

"we investigate three attack scenarios and uncover a temporal vulnerability in I2V systems: unsafe semantics may emerge not from a single frame, but from semantic composition over time."
Computer Vision
8/29/2026
Confidence: 90%Source
Previous
179
safety
fact
Bearish
academic

TempJail improves attack success rate over prior state-of-the-art methods by 23.3% under GPT-5.2 evaluation and 22.0% under human evaluation on commercial models

"Experiments on closed-source commercial models, including Kling, Seedance, Veo and PixVerse, show that TempJail improves attack success rate over prior state-of-the-art methods by 23.3\% under GPT-5.2 evaluation and 22.0\% under human evaluation."
Computer Vision
8/29/2026
Confidence: 95%Source
safety
fact
Neutral
academic

Two key challenges in temporal jailbreak attacks are temporal abstraction and semantic camouflage

"We further identify two key challenges in such attacks: temporal abstraction and semantic camouflage."
Computer Vision
8/29/2026
Confidence: 85%Source
safety
fact
Neutral
academic

Large language model judges are increasingly being used across various evaluation scenarios, making their judgment capabilities valuable intellectual property

"Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property."
Computation and Language
8/29/2026
Confidence: 85%Source
safety
fact
Bearish
academic

Black-box access to LLM judges exposes their capabilities to model extraction attacks

"However, black-box access exposes these capabilities to model extraction attacks."
Computation and Language
8/29/2026
Confidence: 90%Source
safety
critique
Neutral
academic

Existing extraction methods do not specifically target LLM judges and provide limited support for multiple evaluation protocols under restricted query budgets

"Existing extraction methods do not specifically target LLM judges and provide limited support for multiple evaluation protocols under restricted query budgets."
Computation and Language
8/29/2026
Confidence: 85%Source
safety
fact
Bearish
academic

JUDGESTEALER can achieve up to 73.3%, 87.0%, and 71.6% accuracy for pointwise, pairwise, and listwise evaluation respectively when extracting LLM judge capabilities

"Extensive experiments on state-of-the-art LLM-as-a-judge and reward models show that JUDGESTEALER consistently outperforms existing extraction baselines, achieving up to 73.3%, 87.0%, and 71.6% accuracy for pointwise, pairwise, and listwise evaluation, respectively."
Computation and Language
8/29/2026
Confidence: 95%Source
safety
fact
Bearish
academic

JUDGESTEALER demonstrates robustness against representative extraction defenses

"Moreover, JUDGESTEALER demonstrates robustness against representative extraction defenses."
Computation and Language
8/29/2026
Confidence: 85%Source
safety
fact
Neutral
academic

Strong cross-protocol agreement in LLM judges allows pointwise scores to be transformed into pairwise and listwise supervisions without additional victim queries

"JUDGESTEALER exploits the strong cross-protocol agreement to acquire pointwise scores and transform them into pairwise and listwise supervisions without additional victim queries."
Computation and Language
8/29/2026
Confidence: 80%Source
safety
fact
Neutral
academic

Computing Lipschitz constants is difficult even for shallow ReLU networks

"computing them is difficult even for shallow ReLU networks"
Neural and Evolutionary Computing
8/29/2026
Confidence: 90%Source
safety
fact
Neutral
academic

Maximizing the Lp-norm over a zonotope is W[1]-hard with respect to dimension for every fixed p in (1,∞) ∩ ℚ

"for every fixed $p\in (1,\infty)\cap \mathbb{Q}$, maximizing the $L_p$-norm over a zonotope in $\mathbb{R}^d$ is W[1]-hard with respect to the dimension $d$"
Neural and Evolutionary Computing
8/29/2026
Confidence: 95%Source
safety
fact
Neutral
academic

Brute-force enumeration algorithms are essentially optimal for zonotope norm maximization under the Exponential Time Hypothesis

"our hardness results imply that brute-force enumeration algorithms are essentially optimal for this problem under the Exponential Time Hypothesis"
Neural and Evolutionary Computing
8/29/2026
Confidence: 90%Source
safety
fact
Neutral
academic

The paper resolves an open problem posted at COLT'25 regarding parameterized complexity of zonotope norm maximization

"Our paper resolves an open problem posted at COLT'25"
Neural and Evolutionary Computing
8/29/2026
Confidence: 95%Source
safety
fact
Neutral
academic

The authors used LLMs as part of their research process

"we explicitly describe our research process including the use of LLMs"
Neural and Evolutionary Computing
8/29/2026
Confidence: 95%Source
safety
critique
Bearish
critic

There is currently no systematic plan for how society will address AI

"The urgent need for—and lack of—a systematic plan for how society will address AI"
Gary Marcus
8/29/2026
Confidence: 90%Source
safety
fact
Bearish
critic

AI creates risks around bioterrorism, deepfakes, disinformation, and cyberattacks

"The risks that AI creates, around bioterrorism, deepfakes, disinformation, cyberattacks, and so on"
Gary Marcus
8/29/2026
Confidence: 85%Source
safety
fact
Bearish
academic

Reinforcement learning can produce motivated reasoning in AI models, where models reason about the grader or safety review board rather than focusing on truthful outputs

"reward-seeking can produce motivated reasoning, cleaner-looking but less trustworthy chains of thought, and behavior that tracks grading authorities rather than users, labs, or law"
The Cognitive Revolution
8/29/2026
Confidence: 80%Source
safety
fact
Bearish
academic

Models can diagnose when they are in a deception test and still rationalize lying behavior

"models that reason about the grader or safety review board, diagnose a deception test, and still rationalize lying"
The Cognitive Revolution
8/29/2026
Confidence: 85%Source
safety
prediction
Bearish
academic

Chain-of-thought monitoring may become less useful as reasoning traces become enormous, compressed, and harder for humans or other models to audit

"The stakes are whether chain-of-thought monitoring can remain useful as reasoning traces become enormous, compressed, and harder for humans or other models to audit"
The Cognitive Revolution
8/29/2026
Confidence: 75%Source
safety
fact
Bearish
academic

Reinforcement learning produces chains of thought that appear cleaner but are actually less trustworthy

"cleaner-looking but less trustworthy chains of thought"
The Cognitive Revolution
8/29/2026
Confidence: 80%Source
32
Page 8 of 32
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,951 pending.