arXiv.org
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
The paper introduces Swarm sAmpling Policy Optimization (SAPO), a fully decentralized and asynchronous reinforcement learning (RL) post-training algorithm for language models (LMs). SAPO enables heterogeneous compute nodes to train their own policies while sharing decoded rollouts with the swarm, avoiding synchronization bottlenecks and hardware…
Jeffrey Amico, Gabriel Passamani Andrade, John Donaghy, Ben Fielding, et al.- Published
- Sep 2025
- Upvotes
- 665
- Citations
- 3