interpretability
fact
neutral
Moral knowledge geometry in LLMs is consistent across architectures and scale, and emerges early in pre-training before probe accuracy saturates
The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates.
Machine Learning30 Aug 2026