arXiv.org
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
MosaicMem is a hybrid spatial memory mechanism for video world models that combines explicit 3D structure with implicit, attention-based conditioning. It lifts video patches into 3D for precise localization and retrieval, then composes them in the queried view via a patch-and-compose interface, allowing the model to preserve persistent elements while…
Wei Yu, Runjia Qian, Yumeng Li, Liquan Wang, et al.- Published
- Mar 2026
- Citations
- 10
- Code
- Not linked
