Search and filter through extracted claims from AI researchers.
Showing 101-120 of 451 claims in topic "infrastructure"
GLM-5.3-Flash achieves 28% accuracy and 28% hallucination rate on knowledge/factual benchmarks
"AA-Omniscience score: +7Accuracy: 28%Hallucination rate: 28%"
"This suggests a recurring theme in reactions: GLM-5.3-Flash may be much stronger on practical code/agentic workflows than on broad real-world factual knowledge."
GLM-5.3-Flash uses a hybrid attention architecture with 34 KDA layers and 11 MLA/DSA layers
"Kimi Linear-style 3:1 hybrid attention34 KDA layers (Kimi Delta Attention)11 MLA/DSA layersMLA = Multi-head Latent AttentionDSA = DeepSeek Sparse Attention"
GLM-5.3-Flash costs approximately 1/10th of GLM-5.2 while reducing active parameters from 32B to 18B
"compared to GLM-5.2:~1/10 the costactive params 32B → 18Blayers 92 → 45"
"nearly all Chinese frontier models now use linear attentionnearly all use sparse attention / indexer-compression designs"
"Cline said GLM-5.3 Flash was already its fastest growing model in Cline history, driving 11% of all traffic in less than a week"
"says the GLM-5.3 infrastructure agent co-authored parts of the work by helping with kernels, bottlenecks, and serving stack optimization"
"This matters because it reinforces that GLM-5.3-Flash was not treated as a curiosity; it was immediately slotted into real inference/developer stacks."
GLM-5.3-Flash achieves an AA Intelligence Index score of 57 at $0.09 cost per task
"Artificial Analysis reports AA Intelligence Index 57 and $0.09 cost/task, plus various benchmark details and pricing."
GLM-5.3-Flash represents evidence of exciting convergence in Chinese frontier open architectures
"eliebakouch framed it as evidence of exciting convergence in Chinese frontier open architectures."
GLM-5.3-Flash is the best intelligence-per-dollar choice available
"zainhas argued it is now the best intelligence-per-dollar choice."
"In a production-scale A/B test on millions of Spotify users, this policy reduces recommendation impressions by 7% with no statistically significant reduction in overall recommended content consumption."
"argue this can be taken as evidence that causal models learn more generalisable representations than models trained on observational data alone"
"joint training with holdback data improves calibration of the treated head relative to the production baseline"
"extending an existing production recommendation model to a causal architecture using holdback data that is already collected as part of routine experimentation infrastructure, requiring no new data collection"
"Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogonalizing its momentum with a few Newton-Schulz iterations."
"We show that finite Newton-Schulz can instead be beneficial for nonsmooth nonconvex optimization."
"we prove that a Newton-Schulz depth growing only logarithmically in the target accuracy suffices for convergence to stationary points in nonsmooth nonconvex optimization"
Muon with exact-polar update may fail to converge, while Muon with finite Newton-Schulz can succeed
"Muon with the exact-polar update may fail to converge"
"The resulting sample complexity bounds match the best-known guarantees for nonsmooth nonconvex optimization and are optimal for smooth nonconvex optimization up to problem-dependent factors."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.