arXiv.org
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
The paper introduces Spatial Forcing (SF), a method to enhance the spatial awareness of Vision-Language-Action (VLA) models without explicit 3D inputs. VLA models, built on 2D-pretrained VLMs, lack 3D understanding, limiting their robotic manipulation performance. Existing solutions using depth sensors or point clouds face issues like sensor noise and data…
Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, et al.- Published
- Oct 2025
- Citations
- 93
- Code
- 277 stars
