优化薄弱区域,让文档解析模型更准更稳。
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

- 识别并针对性强化模型表现差的区域,提升训练质量。
- 在OmniDocBench上达到96.33%新高分,超越多数顶级视觉语言模型。
- 提供可复用的渐进式后训练方案,适合工程部署与持续优化。
我们提出PaddleOCR-VL-1.6,一个基于PaddleOCR-VL-1.5升级的紧凑型文档解析模型。尽管PaddleOCR-VL-1.5已建立0.9B参数的强基准,但剩余错误集中在模型行为不稳定、数据覆盖稀疏或监督信号不可靠的薄弱区域。与其盲目扩充训练数据,PaddleOCR-VL-1.6引入区域感知的数据优化框架,识别前序模型的弱区域,进行针对性增强,并提升监督信号可靠性。同时采用基于精选数据选择与强化学习的渐进式后训练策略,通过分阶段优化推动性能跃升。该模型在OmniDocBench v1.6上取得96.33%的新纪录,展现出对顶尖视觉语言模型的竞争力,并为PaddleOCR-VL系列提供了实用的后训练方案。
原文摘要 · Abstract (English)
We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining errors concentrate in under-optimized regions where model behavior is unstable, data coverage is sparse, or supervision is unreliable. Rather than expanding the training corpus indiscriminately, PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to these regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score of 96.33% on OmniDocBench v1.6, demonstrates strong competitiveness against top-tier VLMs, and provides a practical post-training recipe for the PaddleOCR-VL series.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。