Search and filter through extracted claims from AI researchers.
Showing 101-120 of 637 claims in topic "safety"
Intent-as-a-tool complements CoT monitoring and expands post-hoc CoT labels into dense trajectories
"Our results show that INTENT-AS-A-TOOL complements CoT monitoring, expands post-hoc CoT labels into dense trajectories, and identifies critical steps for online intervention."
Action preferences are useful for tracking agentic misalignment during reasoning
"These findings suggest that action preferences are useful for tracking agentic misalignment during reasoning."
Adverse drug reactions are a major, largely preventable source of patient harm
"Adverse drug reactions (ADRs) are a major, largely preventable source of patient harm."
"PoP achieves an area under the receiver operating characteristic curve (AUROC) of 75.5% for factual-correctness classification"
"Autoregressive large language models (LLMs) routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows."
Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions
"Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions"
"internal hidden-state transition dynamics during generation can signal factual errors without auxiliary decoding calls"
"The mechanism operates within the base forward pass, adding less than 1.2% runtime latency and requiring zero additional generation passes."
"An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated."
"It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data."
"What unlocks confident action is not information but the authority of its packaging."
"Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy."
"asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails"
"Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains."
"It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly."
"The gate is trainable and context-fragile, and deployment needs both halves of that sentence."
"Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks"
LLM outputs are treated as authoritative even when they are ungrounded or incorrect
"fluent, confident outputs are treated as authoritative even when ungrounded or incorrect"
"of 4,512 records identified, 122 primary studies were included, together with 12 regulatory and standards documents analysed as primary sources"
Current LLM accountability frameworks suffer from under-specification of human oversight
"Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.