HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 101-120 of 637 claims in topic "safety"

safety
fact
Bullish
academic

Intent-as-a-tool complements CoT monitoring and expands post-hoc CoT labels into dense trajectories

"Our results show that INTENT-AS-A-TOOL complements CoT monitoring, expands post-hoc CoT labels into dense trajectories, and identifies critical steps for online intervention."
Computation and Language
8/30/2026
Confidence: 80%Source
Previous
157
safety
opinion
Bullish
academic

Action preferences are useful for tracking agentic misalignment during reasoning

"These findings suggest that action preferences are useful for tracking agentic misalignment during reasoning."
Computation and Language
8/30/2026
Confidence: 75%Source
safety
fact
Neutral
academic

Adverse drug reactions are a major, largely preventable source of patient harm

"Adverse drug reactions (ADRs) are a major, largely preventable source of patient harm."
Machine Learning
8/30/2026
Confidence: 90%Source
safety
fact
Bullish
academic

Prediction of Prediction (PoP) achieves 75.5% AUROC for factual-correctness classification on TruthfulQA

"PoP achieves an area under the receiver operating characteristic curve (AUROC) of 75.5% for factual-correctness classification"
Computation and Language
8/30/2026
Confidence: 90%Source
safety
fact
Bearish
academic

Autoregressive large language models routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows

"Autoregressive large language models (LLMs) routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows."
Computation and Language
8/30/2026
Confidence: 85%Source
safety
critique
Bearish
academic

Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions

"Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions"
Computation and Language
8/30/2026
Confidence: 80%Source
safety
fact
Bullish
academic

Internal hidden-state transition dynamics during generation can signal factual errors without auxiliary decoding calls

"internal hidden-state transition dynamics during generation can signal factual errors without auxiliary decoding calls"
Computation and Language
8/30/2026
Confidence: 80%Source
safety
fact
Bullish
academic

PoP operates within the base forward pass, adding less than 1.2% runtime latency and requiring zero additional generation passes

"The mechanism operates within the base forward pass, adding less than 1.2% runtime latency and requiring zero additional generation passes."
Computation and Language
8/30/2026
Confidence: 90%Source
safety
fact
Bearish
academic

LLM agents commit to directional calls on unpredictable questions 54.0% of the time when shown professional-looking market panels, compared to only 6.5% when asked bare questions, across 12 frontier models

"An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated."
Computation and Language
8/30/2026
Confidence: 90%Source
safety
fact
Bearish
academic

LLMs commit to unpredictable questions just as readily when all data is fabricated, with commitment rising from 24.5% to 36.8% with fake data versus 37.6% with genuine data

"It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data."
Computation and Language
8/30/2026
Confidence: 90%Source
safety
opinion
Bearish
academic

What drives confident LLM action is not information content but the authority of information packaging

"What unlocks confident action is not information but the authority of its packaging."
Computation and Language
8/30/2026
Confidence: 85%Source
safety
fact
Neutral
academic

LLMs can answer matched answerable questions with near-perfect accuracy when attached to the same panels, showing the failure is not due to incapacity

"Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy."
Computation and Language
8/30/2026
Confidence: 85%Source
safety
fact
Bearish
academic

LLMs correctly classify questions as unknowable 90% of the time but then still commit on just 0.4% of those cases, showing the act/don't-act gate is what fails

"asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails"
Computation and Language
8/30/2026
Confidence: 90%Source
safety
fact
Bullish
academic

The act/don't-act gate failure can be fixed through training: supervised fine-tuning of a 3B model on 540 synthetic cases drives commitment on unpredictable questions to 0.0% and transfers to three unseen domains

"Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains."
Computation and Language
8/30/2026
Confidence: 85%Source
safety
fact
Bearish
academic

The trained act/don't-act gate is context-fragile and only holds when response format allows reasoning; rigid formats leave models confident and wrong on questions they otherwise answer correctly

"It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly."
Computation and Language
8/30/2026
Confidence: 85%Source
safety
opinion
Neutral
academic

Deployment of LLMs needs to account for both the trainability of the act/don't-act gate and its context-fragility

"The gate is trainable and context-fragile, and deployment needs both halves of that sentence."
Computation and Language
8/30/2026
Confidence: 80%Source
safety
fact
Neutral
academic

LLMs are currently operating in critical domains including hospitals, courtrooms, banks, and public service desks

"Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
critique
Bearish
academic

LLM outputs are treated as authoritative even when they are ungrounded or incorrect

"fluent, confident outputs are treated as authoritative even when ungrounded or incorrect"
Artificial Intelligence
8/30/2026
Confidence: 80%Source
safety
fact
Bearish
academic

A systematic review of 122 primary studies plus 12 regulatory documents reveals four persistent gaps in LLM accountability frameworks

"of 4,512 records identified, 122 primary studies were included, together with 12 regulatory and standards documents analysed as primary sources"
Artificial Intelligence
8/30/2026
Confidence: 90%Source
safety
critique
Bearish
academic

Current LLM accountability frameworks suffer from under-specification of human oversight

"Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation"
Artificial Intelligence
8/30/2026
Confidence: 85%Source
32
Page 6 of 32
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,951 pending.