Together AI
Beat the Long Tail: Distribution-Aware Speculative Decoding for RL Training
Reinforcement learning (RL) post-training for large language models is bottlenecked by the rollout phase, which accounts for over 70% of training time. The authors identify a long-tail distribution of rollout lengths, where a few long generations dominate wall-clock time, and note that historical rollouts reveal stable prompt-level patterns across epochs.…
Zelei Shao, Vikranth Srivatsa, Sanjana Srivastava, Qingyang Wu, et al.- Published
- Nov 2025
- Citations
- 9
- Code
- Not linked
