The year/Independent research

Paper 2510.18121

Efficient Long-context Language Model Training by Core Attention Disaggregation

Published
Oct 2025
Research lab
Independent
Citations
3
GitHub
Not linked

01 In brief

Summary

This paper introduces core attention disaggregation (CAD), a technique to improve long-context LLM training by separating the parameter-free softmax(QK^T)V computation (core attention, CA) from other model components and scheduling it on a dedicated pool of resources.

CAD leverages two key properties: statelessness (CA has no trainable parameters) and composability (CA can be partitioned into token-level shards and re-batched without losing kernel efficiency).

The system DistCA implements CAD with in-place attention servers (time-sharing GPUs), a ping-pong execution scheme to overlap communication with computation, and a communication-aware greedy scheduler.

Evaluations on up to 512 H200 GPUs with context lengths up to 512K show that DistCA improves end-to-end training throughput by up to 1.35x compared to state-of-the-art systems, eliminates data and pipeline parallelism stragglers, and maintains near-perfect compute and memory balance.

The paper also discusses limitations, such as memory fragmentation overhead, and future work directions including dedicated attention server pools and more flexible sharding strategies.

02 From the paper

Abstract

We present core attention disaggregation (CAD), a technique that improves long-context large language model training by decoupling the core attention computation, softmax(QK^T)V, from the rest of the model and executing it on a separate pool of devices. In existing systems, core attention is colocated with other layers; at long context lengths, its quadratic compute growth compared to the near-linear growth of other components causes load imbalance and stragglers across data and pipeline parallel groups. CAD is enabled by two observations. First, core attention is stateless: it has no trainable parameters and only minimal transient data, so balancing reduces to scheduling compute-bound tasks. Second, it is composable: modern attention kernels retain high efficiency when processing fused batches of token-level shards with arbitrary lengths. CAD partitions core attention into token-level tasks and dispatches them to dedicated attention servers, which dynamically rebatch tasks to equalize compute without sacrificing kernel efficiency. We implement CAD in a system called DistCA, which uses a ping-pong execution scheme to fully overlap communication with computation and in-place execution on attention servers to reduce memory use. On 512 H200 GPUs and context lengths up to 512k tokens, DistCA improves end-to-end training throughput by up to 1.35x, eliminates data and pipeline parallel stragglers, and achieves near-perfect compute and memory balance.