HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 121-140 of 637 claims in topic "safety"

safety
critique
Bearish
academic

There is an absence of shared accountability metrics for LLMs across the field

"absence of shared accountability metrics"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
Previous
168
fact
Bearish
academic

Current agent safeguards defined over single trajectories have a fundamental composition failure: against attacks fragmented across iterations, every trajectory-scoped monitor has equal true-positive and false-positive rates regardless of expressiveness.

"against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
fact
Bearish
academic

Carrying a geometrically decaying risk score is insufficient for agent safety because the cooling-off period an adversary must wait is a constant that doesn't grow with the horizon.

"the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
fact
Bullish
academic

LoopHarness with mediated commits and arbiter detection floor can bound expected unauthorized irreversible actions to a constant independent of horizon N, with part of the bound surviving fully colluding verifiers.

"Under mediated commits and an arbiter detection floor $δ_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/δ_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
fact
Bullish
academic

Large language model agents are increasingly being deployed as autonomous loops that repeatedly discover work, plan, execute, and verify outcomes across many unattended iterations.

"Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations."
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
fact
Neutral
academic

A monitor retaining cross-iteration state can perfectly separate true-positives from false-positives for attacks fragmented across iterations, whereas trajectory-scoped monitors cannot.

"a monitor retaining cross-iteration state separates the two perfectly"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
fact
Bearish
academic

Tool-augmented LLM agents face a security risk when tool outputs specify concrete actions that can drive real-world side effects beyond user intent

"when tool outputs no longer merely provide data but begin to specify concrete actions, they effectively become ``commands'' that can drive real-world side effects beyond user intent"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
opinion
Neutral
academic

The security risk in tool-augmented LLM agents arises from conflating action induction with execution authorization

"this risk arises from conflating action induction with execution authorization"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
fact
Bullish
academic

SARA limits attack success rate to no more than 0.63% across four primary evaluation settings while maintaining competitive task utility

"SARA limits ASR to no more than \(0.63\%\) across four primary evaluation settings while maintaining competitive task utility"
Artificial Intelligence
8/30/2026
Confidence: 95%Source
safety
fact
Bullish
academic

SARA consistently reduces attack success rate across additional agent backbones beyond the primary evaluation

"consistently reduces ASR across additional Agent backbones"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
opinion
Bullish
academic

Separating action induction from execution authorization through distinct runtime roles is an effective approach to securing tool-augmented LLM agents

"we propose SARA, which treats action induction and execution authorization as distinct runtime roles and separates action provenance from execution authority"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
fact
Neutral
academic

The EU AI Act's high-risk obligations have applied since August 2, 2026

"the EU AI Act, whose high-risk obligations have applied since 2 August 2026"
Artificial Intelligence
8/30/2026
Confidence: 95%Source
safety
critique
Bearish
academic

No surveyed accountability instrument resolves five identified structural tensions in LLM governance

"five structural tensions that no surveyed instrument resolves"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
opinion
Neutral
academic

The LAAF integrated accountability architecture is a synthesis of surveyed evidence rather than a validated artifact

"it is a synthesis of the surveyed evidence rather than a validated artefact"
Artificial Intelligence
8/30/2026
Confidence: 95%Source
safety
fact
Neutral
academic

Privacy-Enhancing Technologies in computer vision that rely on noise or image perturbations create a trade-off between task performance and protection

"Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely processing it, creating a trade-off between task performance and protection."
Machine Learning
8/30/2026
Confidence: 90%Source
safety
critique
Neutral
academic

Image classification is too simplistic as a proxy for evaluating privacy-enhancing technologies across generic vision tasks

"This trade-off is commonly evaluated using image classification, which primarily captures semantic separability and remains robust despite significant geometric, spatial layout or local boundary alterations. As a result, it is too simplistic as a proxy for generic vision tasks."
Machine Learning
8/30/2026
Confidence: 85%Source
safety
fact
Neutral
academic

PETs with similar classification accuracy can differ substantially on other vision tasks

"Across irreversible privacy transformations, key-based block primitives, and learnable image encryption schemes, we demonstrate that PETs with similar classification accuracy can differ substantially on other tasks."
Machine Learning
8/30/2026
Confidence: 90%Source
safety
opinion
Neutral
academic

PET evaluation protocols need to move beyond classification-only reporting

"The outcomes highlight the need for PET evaluation protocols that move beyond classification-only reporting."
Machine Learning
8/30/2026
Confidence: 85%Source
safety
fact
Bullish
academic

Image-to-video generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis in recent years

"In recent years, image-to-video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis."
Computer Vision
8/29/2026
Confidence: 90%Source
safety
critique
Bearish
academic

Advances in I2V models introduce new safety risks that existing studies have largely overlooked, particularly in the temporal dimension

"However, these advances also introduce new safety risks. Existing studies mainly focus on jailbreak attacks involving single frame violations, while largely overlooking the temporal dimension unique to video generation models."
Computer Vision
8/29/2026
Confidence: 85%Source
32
Page 7 of 32
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,951 pending.