infrastructure
fact
bullish
ClusterAttention achieves the same latency per query-key interaction as dense attention on GPUs by using fixed power-of-two cluster sizes
We utilize this by setting all clusters to be a fixed size that is a power of two, allowing the block-sparse attention to run at the same latency per query-key interaction as dense attention on GPUs.
Computer Vision29 Aug 2026