Qwen
Qwen3-VL Technical Report
Qwen3-VL is a state-of-the-art vision-language model family from the Qwen team, released on December 1, 2025. It supports interleaved contexts up to 256K tokens and comes in dense (2B/4B/8B/32B) and MoE (30B-A3B/235B-A22B) variants. Key architectural innovations include interleaved-MRoPE for balanced spatial-temporal encoding, DeepStack for multi-level ViT…
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, et al.- Citations
- 1.8K
- Published
- Nov 2025
- Code
- 20K stars