HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 541-560 of 637 claims in topic "safety"

safety
fact
Bearish
academic

Task-specific fine-tuning of LLMs is notoriously plagued by overconfidence, severely hindering trustworthy deployment

Machine Learning
7/27/2026
Confidence: 85%Source
safety
Previous
12729
fact
Bullish
academic

A model-agnostic framework can jointly address privacy preservation and malicious behavior across both federated and decentralized learning settings

Machine Learning
7/27/2026
Confidence: 70%Source
safety
critique
Bearish
academic

Existing defenses for distributed learning typically address privacy and manipulation threats in isolation and are limited to specific paradigms or architectures

Machine Learning
7/27/2026
Confidence: 80%Source
safety
fact
Neutral
academic

Distributed machine learning exposes learning processes to privacy leakage and malicious manipulation

Machine Learning
7/27/2026
Confidence: 85%Source
safety
fact
Bullish
academic

DALorRA provides excellent calibration of LLMs without compromising reasoning accuracy

Machine Learning
7/27/2026
Confidence: 75%Source
safety
fact
Bullish
academic

HaloGuard 1.0-0.8B attains the best average F1 score of 90.9 across seven prompt-safety benchmarks, outperforming baselines up to 27B parameters

Computation and Language
7/27/2026
Confidence: 90%Source
safety
fact
Bullish
academic

HaloGuard 1.0-0.8B achieves state-of-the-art performance on English and multilingual prompt-safety benchmarks at roughly one-tenth the model size of current leading open guard models

Computation and Language
7/27/2026
Confidence: 90%Source
safety
opinion
Neutral
academic

Safe completion should be evaluated as intent-calibrated behavior over controlled task variants, not as aggregate prompt-safety scores

Computation and Language
7/27/2026
Confidence: 85%Source
safety
critique
Bearish
academic

Prompt-level safety evaluation hides important failures in models, as they often fail to remain safe across matched intent variants

Computation and Language
7/27/2026
Confidence: 85%Source
safety
fact
Neutral
academic

Different open instruction-tuned models show distinct policies under skeptical pressure on consensus science topics, with Llama showing reactive assertion rather than false balance

Computation and Language
7/27/2026
Confidence: 90%Source
safety
fact
Bullish
academic

LLMs do not sycophantically retreat from established scientific consensus when users signal doubt, showing reactive assertion, surface hedging, or non-response instead

Computation and Language
7/27/2026
Confidence: 85%Source
safety
critique
Bearish
journalist

Biology and chemistry classifiers in Claude Fable 5 are still overly broad after the relaunch

swyx & Alessio
7/27/2026
Confidence: 70%Source
safety
fact
Bearish
academic

Safety training for LLMs is conducted predominantly in English, creating vulnerabilities in low-resource languages and code-switching scenarios

Computation and Language
7/27/2026
Confidence: 90%Source
safety
fact
Bearish
academic

STEER attack achieves up to 93% success rate on JailbreakBench and 96.7% on AdvBench against 8B-parameter models by exploiting low-resource language gaps

Computation and Language
7/27/2026
Confidence: 95%Source
safety
fact
Bearish
academic

Jailbreak prompts using low-resource language code-switching transfer to GPT-4o-mini with 35.5% success rate without refinement

Computation and Language
7/27/2026
Confidence: 90%Source
safety
fact
Neutral
journalist

Anthropic's Fable 5 is being redeployed with usage limits: included for up to 50% of weekly usage through July 7

Alberto Romero
7/27/2026
Confidence: 95%Source
safety
opinion
Neutral
journalist

The fine print of Fable 5's redeployment reveals important signals about where the AI industry is heading

Alberto Romero
7/27/2026
Confidence: 70%Source
safety
fact
Neutral
academic

In high-noise privacy regimes (strong privacy), increasing privacy reduces generalization error in Byzantine-robust distributed learning, eliminating the tension between robustness and privacy

Machine Learning (Statistics)
7/27/2026
Confidence: 90%Source
safety
fact
Neutral
academic

In low-noise privacy regimes (weaker privacy), the tension between robustness and privacy reappears and increasing privacy degrades generalization in distributed learning

Machine Learning (Statistics)
7/27/2026
Confidence: 90%Source
safety
fact
Neutral
academic

The trilemma between Byzantine robustness, local differential privacy, and optimization error does not universally extend to generalization error

Machine Learning (Statistics)
7/27/2026
Confidence: 85%Source
32
Page 28 of 32
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,951 pending.