The year/Independent research

Paper 2604.04771

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

Published
Apr 2026
Research lab
Independent
Citations
14
GitHub
Not linked

01 In brief

Summary

MinerU2.5-Pro improves document parsing purely through data engineering and training strategy, keeping the 1.2B-parameter architecture of MinerU2.5 unchanged.

The authors identify that state-of-the-art models share failure patterns on hard samples, indicating a data bottleneck rather than an architectural one.

They build a Data Engine with three components: Diversity-and-Difficulty-Aware Sampling (DDAS) expands training data from under 10M to 65.5M pages while balancing distribution; Cross-Model Consistency Verification (CMCV) uses consensus among three heterogeneous models to classify samples into Easy/Medium/Hard tiers; and a Judge-and-Refine pipeline improves annotation quality for hard samples via render-then-verify iterative correction, with residual cases routed to expert annotation.

A three-stage progressive training strategy (large-scale pre-training, hard-sample fine-tuning, and GRPO alignment) leverages these data tiers.

They also introduce OmniDocBench v1.6, which corrects element-matching biases via Multi-Granularity Adaptive Matching and adds a Hard subset for a Base/Hard/Full evaluation protocol.

MinerU2.5-Pro achieves 95.69 on OmniDocBench v1.6, improving over the baseline by 2.71 points and surpassing all existing methods, including models with over 200× more parameters.

The work demonstrates that data-centric optimization can yield significant gains without architectural changes.

02 From the paper

Abstract

Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art models spanning diverse architectures and parameter scales exhibit highly consistent failure patterns on the same set of hard samples, suggesting that the performance bottleneck stems from shared deficiencies in training data rather than from architectural differences. Building on this finding, we present MinerU2.5-Pro, which advances the state of the art purely through data engineering and training strategy design while retaining the 1.2B-parameter architecture of MinerU2.5 unchanged. At its core is a Data Engine co-designed around coverage, informativeness, and annotation accuracy: Diversity-and-Difficulty-Aware Sampling expands training data from under 10M to 65.5M samples while mitigating distribution shift; Cross-Model Consistency Verification leverages output consensus among heterogeneous models to assess sample difficulty and generate reliable annotations; the Judge-and-Refine pipeline improves annotation quality for hard samples through render-then-verify iterative correction. A three-stage progressive training strategy--large-scale pre-training, hard sample fine-tuning, and GRPO alignment--sequentially exploits these data at different quality tiers. On the evaluation front, we rectify element-matching biases in OmniDocBench v1.5 and introduce a Hard subset, establishing the more discriminative OmniDocBench v1.6 protocol. Without any architectural modification, MinerU2.5-Pro achieves 95.69 on OmniDocBench v1.6, improving over the same-architecture baseline by 2.71 points and surpassing all existing methods, including those based on models with over 200x more parameters.