Search and filter through extracted claims from AI researchers.
Showing 61-80 of 171 claims in topic "interpretability"
"we propose an improved method named HiRA-CAM, and show that it outperforms both LayerCAM and Grad-CAM on creating useful saliency maps for object classification."
"We found that input-output alignment was reduced during unconsciousness whereas potential capacity was increased."
DiffusionGemma maintains high monitorability despite having opaque serial depth
"Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance."
"We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be maintained with only the top item, supporting the case for high monitorability."
"we also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. We find that performance is largely retained."
"This is a positive update on the interpretability of diffusion models that are derived from text-pretrained LLMs (an efficient training method more likely to be deployed), but might not apply for more general paradigms."
"Global workspace theory explains conscious access as the broadcasting of selected information to the rest of the network, but it lacks a formal criterion for identifying the mechanism that enables this access."
R-lens directions for intermediate variables are more causally important than J-lens directions
"R-lens directions for intermediate variables are more causally important than J-lens directions"
"We introduce the R-lens: a drop-in replacement for J-lens that produces clearer readouts on earlier layers"
J-lens readouts in early layers are often noisy and largely uninterpretable
"we find readouts in early layers to often be noisy and largely uninterpretable"
R-lens shows a substantial quantitative advantage over J-lens that increases as models scale
"R-lens shows a substantial quantitative advantage over J-lens that increases as models scale, measured across a variety of evaluation categories"
"R-lens produces qualitatively different readouts at earlier layers, significantly reducing the amount of incoherent tokens compared to J-lens"
"R-lens can sometimes capture concepts that J-lens never does, especially if these concepts appear exclusively in early layers"
"Why can deep networks discover abstractions that shallow models miss? Statistical physicist Matthieu Wyart joins Tim Scarfe to argue that the answer lies in the hidden hierarchy of data. Language and images are built from parts within parts; depth lets a network recover those coarse-grained variables and escape the curse of dimensionality."
"why predicting latent representations rather than raw tokens could make learning far more sample-efficient."
The safety community is undervaluing mechanistic interpretability work for alignment
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.