arXiv:2604.04771cs.CVcs.CL2026-04被引 22

通过数据工程提升文档解析性能,无需改动模型架构。

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

  • 设计数据引擎,从1000万增至6550万样本,增强多样性与难度覆盖。
  • 利用多模型一致性验证与迭代修正,提升难例标注质量。
  • 三阶段训练策略有效利用不同质量数据,性能超越大模型。

当前文档解析方法主要依赖模型架构创新,而训练数据的系统性工程仍被忽视。然而,不同架构和参数规模的先进模型在相同难例上表现出高度一致的失败模式,表明性能瓶颈源于共享的训练数据缺陷,而非架构差异。基于此,我们提出MinerU2.5-Pro,仅通过数据工程与训练策略设计,在保持1.2B参数架构不变的前提下实现性能突破。核心为协同设计的数据引擎:多样性和难度感知采样将训练样本从不足1000万扩充至6550万,缓解分布偏移;跨模型一致性验证利用异构模型输出共识评估样本难度并生成可靠标注;判别-优化流水线通过渲染-验证迭代修正难例标注质量。采用三阶段渐进式训练策略——大规模预训练、难例微调与GRPO对齐,分层利用不同质量数据。评估方面,我们修正OmniDocBench v1.5中的元素匹配偏差,引入难例子集,建立更具区分度的OmniDocBench v1.6协议。无需任何架构修改,MinerU2.5-Pro在OmniDocBench v1.6上达到95.69分,相比同架构基线提升2.71分,超越所有现有方法,包括参数量超其200倍的模型。

原文摘要 · Abstract (English)

Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art models spanning diverse architectures and parameter scales exhibit highly consistent failure patterns on the same set of hard samples, suggesting that the performance bottleneck stems from shared deficiencies in training data rather than from architectural differences. Building on this finding, we present MinerU2.5-Pro, which advances the state of the art purely through data engineering and training strategy design while retaining the 1.2B-parameter architecture of MinerU2.5 unchanged. At its core is a Data Engine co-designed around coverage, informativeness, and annotation accuracy: Diversity-and-Difficulty-Aware Sampling expands training data from under 10M to 65.5M samples while mitigating distribution shift; Cross-Model Consistency Verification leverages output consensus among heterogeneous models to assess sample difficulty and generate reliable annotations; the Judge-and-Refine pipeline improves annotation quality for hard samples through render-then-verify iterative correction. A three-stage progressive training strategy--large-scale pre-training, hard sample fine-tuning, and GRPO alignment--sequentially exploits these data at different quality tiers. On the evaluation front, we rectify element-matching biases in OmniDocBench v1.5 and introduce a Hard subset, establishing the more discriminative OmniDocBench v1.6 protocol. Without any architectural modification, MinerU2.5-Pro achieves 95.69 on OmniDocBench v1.6, improving over the same-architecture baseline by 2.71 points and surpassing all existing methods, including those based on models with over 200x more parameters.

文档解析数据工程训练策略模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。