interpretability
fact
neutral
LLMs represent moral tension itself in dilemmas rather than pre-resolved judgments, with dilemma directions partially composing from component foundations but majority variance encoding conflict-specific structure
Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.
Machine Learning30 Aug 2026