arXiv.org
BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
The paper introduces BandPO, a method for LLM reinforcement learning that replaces the fixed clipping bounds of PPO/GRPO with dynamic, probability-aware bounds derived from f-divergence trust regions. The authors identify a bottleneck in canonical clipping: fixed bounds limit the upward update margin for low-probability actions, suppressing high-advantage…
Yuan Li, Bo Wang, Yufei Gao, Yuqian Yao, et al.- Published
- Mar 2026
- Citations
- 2
- Code
- 49 stars
