interpretability
fact
neutral
In CODI-GPT2 and Sim-CoT-style GPT-2 models, counterfactual arithmetic transfers primarily through value-cache suffix trajectories rather than through hidden states, keys, reusable answer slots, or single-token triggers.
On CODI-GPT2 and a Sim-CoT-style GPT-2 reproduction, counterfactual arithmetic transfers primarily through value-cache suffix trajectories rather than hidden states, keys, reusable answer slots, or single-token triggers.
Computation and Language30 Aug 2026