Search and filter through extracted claims from AI researchers.
Showing 1-20 of 171 claims in topic "interpretability"
"Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state."
"CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations"
"On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines"
"while producing a feature-level audit trail of the clinical concepts that support each prediction and the artifact concepts suppressed during training"
"We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component."
"The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013)."
"The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates."
"The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics."
"Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment."
"Visual Retrieval Heads (VRHs), a small subset of attention heads (about 1.7-2.6%) that are causally responsible for grounding text descriptions to image regions"
"Across eleven VLMs and five referring-expression benchmarks, masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, while masking the same number of random heads has little effect"
"they generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction"
VRHs are functionally specific, preserving output format while corrupting localization
"they are functionally specific, preserving output format while corrupting localization"
"they are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning"
"is frequently associated with heightened emotional arousal and distortive communicative intent"
"reader-level interpretations can often be recovered from such limited representations"
"whereas identifying how misleadingness is produced remains considerably more challenging"
"reliable understanding of misleading mechanisms continues to require richer contextual and evidential grounding"
"Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information is framed, omitted, contextualised, and communicated."
Less research has focused on how misleadingness arises and shapes reader interpretations
"Yet less research has focused on how such misleadingness arises and shapes the interpretations formed by readers."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.