arXiv.org
FlowRL: Matching Reward Distributions for LLM Reasoning
FlowRL is a policy optimization algorithm for large language model (LLM) reasoning that shifts from reward maximization to reward distribution matching. It uses a learnable partition function to normalize scalar rewards into a target distribution and minimizes the reverse KL divergence between the policy and this distribution, which is equivalent to a…
Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li, et al.- Published
- Sep 2025
- Citations
- 37
- Code
- Not linked
