Cylon: Asynchronous Linear Attention
Alexander Waitz ⋅ William Hu ⋅ Benjamin F Spector ⋅ Atri Rudra ⋅ Christopher Re ⋅ Simran Arora
Abstract
Sequence models face stark tradeoffs between recall quality and memory efficiency. The ability to use information over long sequences (i.e. recall) is critical for sequence modeling tasks ranging from information extraction to reasoning. Prior work has shown that in theory, *linear* attention models with sufficient recurrent state sizes can expand the Pareto frontier of the recall-memory tradeoff space beyond alternative architectures such as softmax attention and state space models. However, it is difficult to scale the linear attention state size due to hardware bottlenecks. I/O aware algorithms store the linear attention states in thread registers, however state sizes beyond $\approx 3$ megabytes exhaust register memory and trigger expensive register spills. In this work, we introduce Cylon, a hardware-aware strategy for partitioning linear attention's recurrent state across the registers of multiple GPU processors and asynchronously combining the partitions. When applying Cylon to popular architectures, such as Hedgehog, Mamba-2, and Gated Deltanet, we unlock $6 \times$ higher throughput compared to prior linear attention algorithms for these architectures on both Hopper and Blackwell GPUs. Finally, Cylon makes large states available to model designers by unlocking sizes (e.g. $\geq 131$ MB) that are not achievable by the existing linear attention kernels.
Successful Page Load