Search and filter through extracted claims from AI researchers.
Showing 1-16 of 16 claims in topic "interpretability" of type "critique"
"parallel analysis-derived component counts and decisions can reflect hidden-coordinate choice rather than a well-defined property of the model"
"Because each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them."
"existing post-hoc methods often ignore temporal dependence and fail to provide horizon-specific explanations"
"Global workspace theory explains conscious access as the broadcasting of selected information to the rest of the network, but it lacks a formal criterion for identifying the mechanism that enables this access."
J-lens readouts in early layers are often noisy and largely uninterpretable
"we find readouts in early layers to often be noisy and largely uninterpretable"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.