HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 1-20 of 171 claims in topic "interpretability"

interpretability
fact
Bearish
academic

Clinical language models exploit note-specific artifacts (templates, separators, boilerplate) that do not reflect patient state, causing them to fail under deployment shifts despite strong in-hospital accuracy

"Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state."
Computation and Language
8/30/2026
Confidence: 85%Source
29
Page 1 of 9Next
interpretability
fact
Bullish
academic

CAST (Concept-guided Artifact Suppression Tuning) uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations for clinical text classification

"CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations"
Computation and Language
8/30/2026
Confidence: 90%Source
interpretability
fact
Bullish
academic

CAST improves over fine-tuned encoder baselines and remains competitive with strong LLM baselines on MIMIC-IV discharge-note mortality prediction

"On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines"
Computation and Language
8/30/2026
Confidence: 85%Source
interpretability
fact
Bullish
academic

SAE-based approaches can provide auditable, feature-level audit trails showing clinical concepts supporting predictions and artifact concepts suppressed during training

"while producing a feature-level audit trail of the clinical concepts that support each prediction and the artifact concepts suppressed during training"
Computation and Language
8/30/2026
Confidence: 90%Source
interpretability
fact
Neutral
academic

Large language models organize moral knowledge geometrically, with moral foundation representations spanning near-maximal independent dimensions while sharing a positive common component

"We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component."
Machine Learning
8/30/2026
Confidence: 85%Source
interpretability
fact
Neutral
academic

The shared component in moral representation is moral-specific with much higher integration compared to matched non-moral concept batteries

"The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013)."
Machine Learning
8/30/2026
Confidence: 90%Source
interpretability
fact
Neutral
academic

Moral knowledge geometry in LLMs is consistent across architectures and scale, and emerges early in pre-training before probe accuracy saturates

"The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates."
Machine Learning
8/30/2026
Confidence: 85%Source
interpretability
fact
Neutral
academic

LLM moral representations reflect corpus statistics rather than the individualizing/binding distinction predicted by Moral Foundations Theory

"The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics."
Machine Learning
8/30/2026
Confidence: 75%Source
interpretability
fact
Neutral
academic

LLMs represent moral tension itself in dilemmas rather than pre-resolved judgments, with dilemma directions partially composing from component foundations but majority variance encoding conflict-specific structure

"Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment."
Machine Learning
8/30/2026
Confidence: 80%Source
interpretability
fact
Neutral
academic

Vision-language models contain Visual Retrieval Heads (VRHs), a small subset of about 1.7-2.6% of attention heads that are causally responsible for grounding text descriptions to image regions

"Visual Retrieval Heads (VRHs), a small subset of attention heads (about 1.7-2.6%) that are causally responsible for grounding text descriptions to image regions"
Computer Vision
8/30/2026
Confidence: 90%Source
interpretability
fact
Neutral
academic

Masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points across eleven VLMs and five benchmarks, demonstrating their causal importance

"Across eleven VLMs and five referring-expression benchmarks, masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, while masking the same number of random heads has little effect"
Computer Vision
8/30/2026
Confidence: 95%Source
interpretability
fact
Bullish
academic

Visual Retrieval Heads generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction

"they generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction"
Computer Vision
8/30/2026
Confidence: 85%Source
interpretability
fact
Neutral
academic

VRHs are functionally specific, preserving output format while corrupting localization

"they are functionally specific, preserving output format while corrupting localization"
Computer Vision
8/30/2026
Confidence: 85%Source
interpretability
fact
Bullish
academic

VRHs are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning

"they are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning"
Computer Vision
8/30/2026
Confidence: 85%Source
interpretability
fact
Neutral
academic

Misleadingness is frequently associated with heightened emotional arousal and distortive communicative intent

"is frequently associated with heightened emotional arousal and distortive communicative intent"
Computation and Language
8/30/2026
Confidence: 85%Source
interpretability
fact
Bullish
academic

Reader-level interpretations can often be recovered from lightweight claim-and-context representations without access to richer contextual, evidential, and multimodal information

"reader-level interpretations can often be recovered from such limited representations"
Computation and Language
8/30/2026
Confidence: 85%Source
interpretability
fact
Bearish
academic

Identifying how misleadingness is produced from lightweight representations remains considerably more challenging than recovering reader interpretations

"whereas identifying how misleadingness is produced remains considerably more challenging"
Computation and Language
8/30/2026
Confidence: 90%Source
interpretability
opinion
Neutral
academic

Reliable understanding of misleading mechanisms requires richer contextual and evidential grounding beyond lightweight representations

"reliable understanding of misleading mechanisms continues to require richer contextual and evidential grounding"
Computation and Language
8/30/2026
Confidence: 80%Source
interpretability
fact
Neutral
academic

Influential public discourse can mislead not only through what is stated, but also through how information is framed, omitted, contextualized, and communicated

"Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information is framed, omitted, contextualised, and communicated."
Computation and Language
8/30/2026
Confidence: 90%Source
interpretability
fact
Neutral
academic

Less research has focused on how misleadingness arises and shapes reader interpretations

"Yet less research has focused on how such misleadingness arises and shapes the interpretations formed by readers."
Computation and Language
8/30/2026
Confidence: 80%Source

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,947 pending.