Search and filter through extracted claims from AI researchers.
Showing 41-60 of 171 claims in topic "interpretability"
"across five models, three retrieval domains, and 75 transformations, median component-count disagreement is 0.79 and median fixed-threshold decision disagreement is 0.26"
"parallel analysis-derived component counts and decisions can reflect hidden-coordinate choice rather than a well-defined property of the model"
"Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning."
"Because each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them."
"On synthetic data with known ground truth, ICON recovers concept importance more accurately than seven alternative baseline methods."
"On skin-lesion and brain-imaging models, it isolates the concepts on which a model genuinely relies, quantifies the representation unexplained by any of the supplied concepts, and yields sparse explanations that we validate by retraining and out-of-distribution testing."
"existing post-hoc methods often ignore temporal dependence and fail to provide horizon-specific explanations"
"We propose a model-agnostic explainability framework that explains forecasting predictions by attributing each forecast horizon to temporally relevant historical lags"
"the semantic-flow variant achieves competitive or superior faithfulness compared to standard post-hoc baselines, while being substantially more computationally efficient"
"Stability analysis further demonstrates that the explanations are robust and identifies regimes where interpretation should be applied with caution"
Goodfire has built Silico, a research platform priced at $1,000 per month
"Silico, the $1,000-per-month research platform Goodfire built for itself"
"fine-tuning and RL often amplify behaviors already latent in pre-training"
Interpretability techniques can identify the data and features driving unwanted model updates
"interpretability can identify the data and features driving unwanted updates"
"models do not store concepts as simple one-hot features, but as sparse mixtures of meaningful subspaces whose geometry determines what kinds of steering and control work"
"sparse mixtures of meaningful subspaces whose geometry determines what kinds of steering and control work"
Modern interpretability research may be moving beyond its reputation as a toy-model science
"modern interpretability may be moving beyond its reputation as a toy-model science"
Steering techniques can fail when applied off-manifold from the concept subspaces
"why steering can fail off-manifold"
"Closing the path, by taking the same statistics per position, multiplies what enlarging the receptive field is worth by up to an order of magnitude on simulated genomes at every difficulty level tested and on real 1000 Genomes haplotypes."
"This graphical notation gives a global view of an architecture and maps one to one onto PyTorch einsum code."
The diagram of Steerling-8B architecture translates into just 33 lines of PyTorch code
"The diagram yields global insights into the architecture (e.g., showing that Steerling is a residual model), a geometric interpretation of each individual operation, and a direct translation into 33 lines of PyTorch code."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.