arXiv.org
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
Vision-Zero is a label-free, domain-agnostic multi-agent self-play framework for self-evolving vision-language models (VLMs) through competitive visual games generated from arbitrary images. It trains VLMs in a 'Who Is the Spy?'-style game where civilians see an image and the spy sees a blank input, requiring strategic reasoning and communication. The…
Qinsi Wang, Bo Liu, Tianyi Zhou, Jing Shi, et al.- Published
- Sep 2025
- Citations
- 33
- Code
- 165 stars
