Paper 2602.06932
When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
- Published
- Feb 2026
- Research lab
- Together AI
- Citations
- 1
- GitHub
- Not linked
01 In brief
Summary
Aurora is a unified training-serving system that addresses limitations of conventional speculative decoding, which separates offline speculator training from online serving, causing deployment lag, delayed utility feedback, and domain-drift degradation.
Aurora closes the loop by continuously learning a speculator from live inference traces, framing it as an asynchronous reinforcement learning problem where accepted tokens provide positive feedback and rejected proposals provide negative feedback.
The system integrates an SGLang-based inference server with an asynchronous training server, enabling hot-swapped speculator updates without service interruption and supporting day-0 deployment from scratch.
Experiments show Aurora achieves a 1.5× day-0 speedup on frontier models like MiniMax M2.1 229B and Qwen3-Coder-Next 80B, and adapts to distribution shifts with an additional 1.25× speedup over static speculators on Qwen3 and Llama3.
Key findings include that simple online fine-tuning captures most gains, lazy synchronization balances adaptation speed with serving stability, and training on discarded tokens provides marginal benefits except with longer lookahead.
Aurora also reduces infrastructure costs by eliminating large-scale activation collection and replay pipelines, and is algorithm-agnostic, supporting various speculative decoding variants.
The system scales to large GPU deployments and heterogeneous request mixtures, demonstrating robust performance across batch sizes and model architectures.
02 From the paper
Abstract
Speculative decoding can significantly accelerate LLM serving, yet most deployments today disentangle speculator training from serving, treating speculator training as a standalone offline modeling problem. We show that this decoupled formulation introduces substantial deployment and adaptation lag: (1) high time-to-serve, since a speculator must be trained offline for a considerable period before deployment; (2) delayed utility feedback, since the true end-to-end decoding speedup is only known after training and cannot be inferred reliably from acceptance rate alone due to model-architecture and system-level overheads; and (3) domain-drift degradation, as the target model is repurposed to new domains and the speculator becomes stale and less effective. To address these issues, we present Aurora, a unified training-serving system that closes the loop by continuously learning a speculator directly from live inference traces. Aurora reframes online speculator learning as an asynchronous reinforcement-learning problem: accepted tokens provide positive feedback, while rejected speculator proposals provide implicit negative feedback that we exploit to improve sample efficiency. Our design integrates an SGLang-based inference server with an asynchronous training server, enabling hot-swapped speculator updates without service interruption. Crucially, Aurora supports day-0 deployment: a speculator can be served immediately and rapidly adapted to live traffic, improving system performance while providing immediate utility feedback. Across experiments, Aurora achieves a 1.5x day-0 speedup on recently released frontier models (e.g., MiniMax M2.1 229B and Qwen3-Coder-Next 80B). Aurora also adapts effectively to distribution shifts in user traffic, delivering an additional 1.25x speedup over a well-trained but static speculator on widely used models (e.g., Qwen3 and Llama3).