arXiv:2510.19817cs.CVcs.CL2025-10被引 42

用测试奖励训练的70亿参数模型,让文档识别更准更稳。

olmOCR 2: Unit Test Rewards for Document OCR

  • 用可验证的单元测试做奖励,通过强化学习训练模型。
  • 在数学公式、表格和多栏布局上表现显著优于前代。
  • 适合需要高精度文档转换的研究者与开发者使用。

我们提出 olmOCR 2,一款用于将数字化印刷文档(如 PDF)转化为清晰、自然排序纯文本的强大 OCR 系统。olmOCR 2 基于 olmOCR-2-7B-1025 模型,一个 70 亿参数的专用视觉语言模型,采用基于可验证奖励的强化学习(RLVR)进行训练,其奖励来源于多样化的二元单元测试。为规模化生成测试用例,我们构建了合成文档生成流水线,涵盖复杂多样的版式布局、已知真实 HTML 源码及提取的测试用例。实验表明,基于这些测试用例的强化学习训练,在我们的英语 OCR 基准 olmOCR-Bench 上达到当前最优性能,尤其在数学公式转换、表格解析和多列布局处理方面提升显著。模型、数据与代码已开源,采用宽松许可协议。

原文摘要 · Abstract (English)

We present olmOCR 2, the latest in our family of powerful OCR systems for converting digitized print documents, like PDFs, into clean, naturally ordered plain text. olmOCR 2 is powered by olmOCR-2-7B-1025, a specialized, 7B vision language model (VLM) trained using reinforcement learning with verifiable rewards (RLVR), where our rewards are a diverse set of binary unit tests. To scale unit test creation, we develop a pipeline for generating synthetic documents with diverse and challenging layouts, known ground-truth HTML source code, and extracted test cases. We show that RL training on these test cases results in state-of-the-art performance on olmOCR-Bench, our English-language OCR benchmark, with the largest improvements in math formula conversion, table parsing, and multi-column layouts compared to previous versions. We release our model, data and code under permissive open licenses.

OCR强化学习文档理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。