interpretability
fact
neutral
Large language models organize moral knowledge geometrically, with moral foundation representations spanning near-maximal independent dimensions while sharing a positive common component
We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component.
Machine Learning30 Aug 2026