arXiv.org
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
MinerU2.5 is a 1.2B-parameter vision-language model for document parsing that achieves state-of-the-art accuracy with high efficiency. It uses a decoupled, two-stage strategy: first, global layout analysis on a downsampled 1036x1036 image; second, targeted content recognition on native-resolution crops guided by the layout. The model uses a 675M NaViT…
Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang, et al.- GitHub stars
- 77K
- Citations
- 86
- Published
- Sep 2025