interpretability
fact
neutral
LLM moral representations reflect corpus statistics rather than the individualizing/binding distinction predicted by Moral Foundations Theory
The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics.
Machine Learning30 Aug 2026