Research paper
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
The paper introduces Zone of Proximal Policy Optimization (ZPPO), a post-training method for small vision-language models (VLMs) that transfers knowledge from a larger teacher without imitating its logits or injecting its responses into the policy gradient. ZPPO addresses two failure modes: distillation's brittleness in the small-student regime and RL's…
Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang, et al.- Published
- Jun 2026
- Citations
- 2
- Code
- Not linked