arXiv.org
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Tuna-2 is a native unified multimodal model that performs visual understanding and generation directly from raw pixel embeddings, eliminating pretrained vision encoders such as VAEs and representation encoders. It uses simple patch embedding layers to encode images and a single transformer decoder for joint processing, with pixel-space flow matching for…
Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen, et al.- Published
- Apr 2026
- Citations
- 10
- Code
- 744 stars