Paper 2607.05394
Weak-to-Strong Generalization via Direct On-Policy Distillation
- Published
- Jul 2026
- Research lab
- Independent
- Citations
- 4
- GitHub
- Not linked
01 In brief
Summary
The paper introduces Direct On-Policy Distillation (Direct-OPD), a method to transfer the policy shift induced by reinforcement learning (RL) on a small, weak teacher model to a stronger student model, avoiding the high cost of running RL directly on the larger model.
Instead of imitating the post-RL teacher's final policy, Direct-OPD uses the log-ratio between the post-RL teacher and its pre-RL reference as a dense implicit reward, applied on the student's own on-policy states.
This signal is mathematically equivalent to the teacher's reward under KL-regularized RL.
Experiments show that Direct-OPD consistently improves stronger students, including those already outperforming the teacher, across different teacher pairs and model families.
Notably, it boosts Qwen3-1.7B from 48.3% to 58.3% on AIME 2024 in about 4 hours on 8 A100 GPUs, outperforming step-matched direct RL and enabling sequential composition of multiple policy shifts.
The method does not require high teacher-student token overlap and generalizes beyond the training horizon.
An adaptive KL controller balances the dense reward signal, as the optimal KL strength is pair-dependent.
Limitations include potential failure when the teacher's improvement is not meaningful on student-visited states, and the need to tune response length and KL strength per pair.
02 From the paper
Abstract
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-strong alternative: run RL on a smaller model where rollouts are cheaper, then reuse what that RL run learned to improve a stronger target model. Directly distilling the post-RL weak teacher is not enough, because the teacher's final policy mixes useful RL gains with the limitations of the smaller model. We propose Direct On-Policy Distillation (Direct-OPD), which transfers the teacher's RL-induced policy shift instead. Direct-OPD compares the post-RL teacher with its own pre-RL reference and treats their log-ratio as a dense implicit reward for the student. In plain terms, the checkpoint pair tells us which actions RL made the weak model more or less likely to take, and Direct-OPD applies that signal on the stronger student's own on-policy states. This directly reuses the weak model's RL supervision signal without running sparse-reward RL on the target model. Empirically, Direct-OPD consistently leverages weaker teachers to improve stronger target models; notably, it boosts Qwen3-1.7B from 48.3% to 58.3% on AIME 2024 in just 4 hours on 8 A100 GPUs. It outperforms step-matched direct RL and enables the sequential composition of multiple policy shifts. Our results show that RL outcomes can be reused across model scales as implicit reward signals, not merely as final models to imitate.