Search and filter through extracted claims from AI researchers.
Showing 81-100 of 451 claims in topic "infrastructure"
Attention magnitude is unrelated to a token's causal contribution to the answer in KV cache eviction
"Using a controlled leave-one-out probe, we find that attention magnitude is unrelated to a token's causal contribution to the answer (Spearman $ρ=-0.004$), challenging the premise behind dominant eviction methods."
"On Qwen3-4B, TwinKV improves a majority of configurations for two policies, is near-even for a third, and helps only a minority for a fourth adaptive baseline already near a performance ceiling"
Long-context inference is bottlenecked by KV cache memory footprint, especially for small models
"Long-context inference is bottlenecked by the memory footprint of the key-value (KV) cache, especially for small models under tight resource budgets."
TwinKV does not help performance on few-shot classification exemplar tasks
"We also identify few-shot classification exemplars as a task structure where TwinKV does not help on either model."
OpenAI's Jalapeño chip will start deploying by year-end
"It starts deploying by year-end."
"This paper introduces ClusterAttention, a general training-free speedup of bidirectional attention layers. Existing sparse attention methods either rely on structure in the input, such as order in language or spatial proximity in images, or use slow clustering processes amortized over several forward passes. ClusterAttention instead uses a fast recursive clustering method that adapts to the geometry of the keys and queries in each attention head to produce useful clusters."
"We utilize this by setting all clusters to be a fixed size that is a power of two, allowing the block-sparse attention to run at the same latency per query-key interaction as dense attention on GPUs."
"On large-scale tabular data ClusterAttention speeds up TabPFN-3 arXiv:2605.13986 by two to six times, while retaining at least 99% of the dense accuracy."
"To our knowledge, it is the first training-free method that can be successfully applied in the setting of unstructured input and a single forward pass."
"For video generation with Wan 2.1-14B T2V arXiv:2503.20314 , ClusterAttention achieves output closer to dense attention and a larger speedup (1.8x versus 1.4x) compared to SVOO arXiv:2603.18636 , a leading method developed specifically for this domain, both run without offline calibration."
"We also derive an expression for the output error in sparse attention, that explains the counterintuitive experimental finding that tight clusters can lead to larger errors than random clusters. We then derive the error when excluded clusters are compensated through their centroids, and show that this error shrinks with tighter clusters."
"FoldPipe records 16.33 s mean I/O-compute overlap, compared with zero by construction for the sequential bounded baseline."
"Mean pass time is 76.78 s for FoldPipe and 83.37 s for the baseline. However, the geometric mean paired speedup is $1.059\times$ with a 95% bootstrap interval from $0.878\times$ to $1.288\times$."
"The experiment therefore verifies the overlap mechanism but is inconclusive about a reliable wall-clock speed advantage under the observed public-network variability."
"Asynchronous prefetch and bounded buffering are established systems techniques rather than novel scheduling algorithms. FoldPipe's contribution is a small integration targeted at native .pt molecular shards together with a source-pinned empirical characterization of its operating regime."
"The conventional wisdom in AI is that the next breakthrough will come from more compute, more data, and larger models. But what if the next leap comes from somewhere else?"
Physics may provide ideas behind the next generation of AI systems
"Max Welling—co-founder and CTO of CuspAI and professor at the University of Amsterdam—argues that physics may provide some of the ideas behind the next generation of AI systems."
Waves may become a new computational primitive for neural networks
"why waves may become a new computational primitive for neural networks"
Nvidia is acquiring HuggingFace for $13B, roughly 80x their $150M ARR
"Nvidia is buying HuggingFace for $13B, roughly 80x their $150M ARR, having doubled its customer base in 2026"
HuggingFace doubled its customer base in 2026
"having doubled its customer base in 2026"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.