Search and filter through extracted claims from AI researchers.
Showing 21-40 of 171 claims in topic "interpretability"
"misleadingness is diverse and extends well beyond fabrication, with unsupported inference, exaggeration, and omission among the prevalent mechanisms"
"While deep transformers have been widely studied in computer vision and natural language processing, their application in tabular data remains relatively underexplored."
"In 72.5\% of experimental examples, the model remains most resilient to performance drops when heads with the lowest importance scores are gradually removed."
"removing the most important attention head first results in the greatest reduction in classification performance"
"A closer look at individual head importance scores across six attention layers reveals that important heads are scattered across layers, with no consistent layer-specific trends."
"In contrast to the image and language domains, the importance of individual attention heads varies considerably across tabular datasets with different schemas and feature spaces."
Head importance scoring can improve efficiency and reduce redundancy in transformer architectures
"The proposed importance score can improve efficiency and redundancy within transformer architectures."
"Across four behaviors and eight models, condensed circuits are smaller than the strongest frozen baseline in 30 of 32 settings, by $8.1\times$ on average and up to $316\times$."
"Repeating the search without weight updates produces larger circuits in 29 of 32 settings, showing that weight updates, rather than search alone, drive the reduction."
"On indirect object identification, condensation isolates 24 heads, 17 of them with documented roles, against 61 heads and 36 undocumented ones for the matched frozen circuit: a sufficient sub-circuit of the published mechanism rather than a reconstruction of it."
Edges in condensed circuits have dependencies and their effects cannot be understood independently
"Pair ablations expose dependencies between edges, showing that their effects cannot be understood independently."
"Testing every subset of 19 circuits finds 11 that cannot be reduced and reveals removable edges in the rest."
"Latent chain-of-thought models move intermediate reasoning from emitted text into continuous states, improving compactness but hiding the causal object."
"We introduce SCIT, the Suffix Cache Interchange Test, a causal protocol that constructs exact source-recipient counterfactuals, patches declared cache segments, and identifies which transformer object carries the counterfactual computation."
"On CODI-GPT2 and a Sim-CoT-style GPT-2 reproduction, counterfactual arithmetic transfers primarily through value-cache suffix trajectories rather than hidden states, keys, reusable answer slots, or single-token triggers."
"SCIT reveals carrier-regime shifts: arithmetic-like GPT-2/1B cells preserve latent-tail value/KV transfer, whereas competent 8B and repaired non-arithmetic cells route through prompt-prefix or full-cache K/V"
"SCIT therefore contributes a cache-level diagnostic, a checkpoint-specific GPT-2 arithmetic mechanism, and a competence-gated carrier map rather than a universal latent-tail claim."
"Hidden coordinates are not uniquely determined by a language model's input--output function, so representation-derived measurements should be invariant to function-preserving changes of basis."
"column-permutation parallel analysis violates function-preserving reparameterization invariance because its reference distribution and selected component count can change while the model function and observed covariance spectrum remain fixed"
"a data-internal reference procedure cannot simultaneously preserve every coordinate marginal, remain orthogonally equivariant, and remove cross-coordinate covariance"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.