infrastructure
fact
neutral
GLM-5.3-Flash uses a hybrid attention architecture with 34 KDA layers and 11 MLA/DSA layers
Kimi Linear-style 3:1 hybrid attention34 KDA layers (Kimi Delta Attention)11 MLA/DSA layersMLA = Multi-head Latent AttentionDSA = DeepSeek Sparse Attention
swyx & Alessio29 Aug 2026
https://www.latent.space/p/ainews-nvidia-buys-huggingface-for