HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 81-100 of 637 claims in topic "safety"

safety
fact
Bearish
academic

LLM-based agents deployed in product-level execution harnesses create greater risks through jailbreaks triggering harmful tool use and persistent state changes than unsafe text generation alone

"LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone."
Artificial Intelligence
8/30/2026
Confidence: 85%Source
Previous
146
safety
fact
Neutral
academic

ModelAudit produced definitive security decisions for all 135 labeled artifact families, achieving 100% coverage

"ModelAudit produced definitive security decisions for all 135 families (100%)"
Artificial Intelligence
8/30/2026
Confidence: 95%Source
safety
fact
Bullish
academic

ModelScan achieved perfect precision, recall, and F1 scores when it was able to make definitive judgments

"Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1."
Artificial Intelligence
8/30/2026
Confidence: 95%Source
safety
fact
Bearish
academic

ModelScan only produced definitive security decisions for 49.6% of labeled families, showing limited coverage despite high accuracy

"ModelScan for 67 (49.6%)"
Artificial Intelligence
8/30/2026
Confidence: 95%Source
safety
fact
Neutral
academic

Fickling provided no unique detection value beyond the combination of ModelAudit and ModelScan

"Fickling identified no unique true- positive families beyond those found by the combination of ModelAudit and ModelScan."
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
opinion
Neutral
academic

Evaluation metrics for ML artifact security scanners must distinguish judgment accuracy from judgment availability

"These findings underscore the need to separate judgment accuracy from judgment availability, as well as incremental detection coverage from tool-level redundancy."
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
fact
Neutral
academic

For the 48 malicious families where ModelScan failed analysis, both ModelAudit and Fickling successfully generated accurate detections

"for the 48 malicious families where ModelScan failed to complete its analysis, both ModelAudit and Fickling generated detections consistent with ground truth."
Artificial Intelligence
8/30/2026
Confidence: 95%Source
safety
opinion
Neutral
academic

AI-generated text detection systems should treat mixed-origin writing as a multi-dimensional problem distinguishing content origin from expression origin rather than a simple binary classification

"AI-generated text detection is commonly framed as a binary document-level judgment about whether a text is human-written or machine-generated. This framing breaks down for mixed-origin writing, where content origin and expression origin may differ."
Computation and Language
8/30/2026
Confidence: 85%Source
safety
fact
Neutral
academic

The D2C-Routing detector system achieves 0.8603 four-way Avg TPR@1%FPR on mixed-origin text detection, outperforming RACE-local by 6.5 points

"our disclosed D2C-Routing-based detector system reaches 0.8603 four-way Avg TPR@1%FPR, 6.5 points above the same-split RACE-local rerun"
Computation and Language
8/30/2026
Confidence: 95%Source
safety
fact
Bearish
academic

Distinguishing AI-content/human-expression from fully AI-generated text is the hardest classification boundary in mixed-origin detection

"error analysis shows that distinguishing AI-content/human-expression from fully AI-generated text remains the hardest boundary"
Computation and Language
8/30/2026
Confidence: 80%Source
safety
fact
Neutral
academic

Verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored or safety-flavored, and these framings predict compliance very differently

"verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
fact
Neutral
academic

On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions

"On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
fact
Neutral
academic

A CoT-prefill intervention demonstrates a causal link between eval-awareness framing and compliance, with 10 of 11 prefills shifting compliance in the predicted direction

"A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
safety
fact
Neutral
academic

Eval-awareness is not behaviorally uniform - aggregate suppression rates can move while the safety-relevant component does not

"eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not"
Artificial Intelligence
8/30/2026
Confidence: 80%Source
safety
fact
Bearish
academic

The same percentage of eval-awareness suppression can correspond to qualitatively different behavioral outcomes

"the same "X% suppression of eval-awareness" can correspond to qualitatively different behavioral outcomes"
Artificial Intelligence
8/30/2026
Confidence: 80%Source
safety
critique
Bearish
academic

Current safety evaluation pipelines treat eval-awareness as a single quantity to be suppressed, which may be inadequate given its heterogeneous nature

"Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed"
Artificial Intelligence
8/30/2026
Confidence: 75%Source
safety
fact
Bearish
academic

As LLMs are deployed as autonomous agents, safety failures increasingly involve consequential actions

"As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions."
Computation and Language
8/30/2026
Confidence: 80%Source
safety
fact
Neutral
academic

Harmful execution in agentic systems is often preceded by intent signals in reasoning when using chain-of-thought monitoring

"Using chain-of-thought (CoT) monitoring, we find that harmful execution is often preceded by intent signals in reasoning."
Computation and Language
8/30/2026
Confidence: 85%Source
safety
critique
Neutral
academic

Post-hoc chain-of-thought labels are too coarse to show how intent changes during generation

"However, post-hoc CoT labels are too coarse to show how intent changes during generation."
Computation and Language
8/30/2026
Confidence: 80%Source
safety
fact
Bullish
academic

Intent-as-a-tool approach provides a judge-free, fine-grained signal of the model's tendency to pursue specific behaviors

"The probability of calling an intent tool provides a judge-free, fine-grained signal of the model's tendency to pursue that behavior."
Computation and Language
8/30/2026
Confidence: 85%Source
32
Page 5 of 32
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,951 pending.