arXiv.org
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
Green-VLA is a staged Vision-Language-Action (VLA) framework for real-world robot deployment, developed by Sber Robotics Center. It uses a five-stage curriculum: L0 (base VLM), L1 (web pretraining), R0 (multi-embodiment robotics pretraining), R1 (embodiment-specific fine-tuning), and R2 (RL alignment). The framework unifies 24M web samples and 3,000 hours…
I. Apanasevich, M. Artemyev, R. Babakyan, P. Fedotova, et al.- Published
- Jan 2026
- Upvotes
- 322
- Citations
- 4