171 claims over the last 90 days
No lab researcher claims on this topic.
Clinical language models exploit note-specific artifacts (templates, separators, boilerplate) that do not reflect patient state, causing them to fail under deployment shifts despite strong in-hospital accuracy
CAST (Concept-guided Artifact Suppression Tuning) uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations for clinical text classification
CAST improves over fine-tuned encoder baselines and remains competitive with strong LLM baselines on MIMIC-IV discharge-note mortality prediction
SAE-based approaches can provide auditable, feature-level audit trails showing clinical concepts supporting predictions and artifact concepts suppressed during training
Large language models organize moral knowledge geometrically, with moral foundation representations spanning near-maximal independent dimensions while sharing a positive common component
The shared component in moral representation is moral-specific with much higher integration compared to matched non-moral concept batteries
Moral knowledge geometry in LLMs is consistent across architectures and scale, and emerges early in pre-training before probe accuracy saturates
LLM moral representations reflect corpus statistics rather than the individualizing/binding distinction predicted by Moral Foundations Theory
LLMs represent moral tension itself in dilemmas rather than pre-resolved judgments, with dilemma directions partially composing from component foundations but majority variance encoding conflict-specific structure
Vision-language models contain Visual Retrieval Heads (VRHs), a small subset of about 1.7-2.6% of attention heads that are causally responsible for grounding text descriptions to image regions
No other claims on this topic.