将论文页面还原为可编译的LaTeX代码,提升科学文档的自动化处理质量。
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction

- 构建了面向可编译LaTeX重建的多维度评估基准和大规模训练数据集。
- 20亿参数模型通过强化学习优化,使生成内容在结构与编译上更可靠。
- 适合需要高精度文档数字化的科研人员和出版系统开发者。
现有文档OCR主要针对纯文本或Markdown,忽视了使LaTeX成为科学出版核心的结构与可执行特性。本文研究从科学PDF页面级重建为可编译LaTeX的任务,提出TexOCR-Bench基准和TexOCR-Train大规模训练语料库。TexOCR-Bench包含多维度评估体系,联合评估转录准确率、结构忠实度和端到端编译性。基于TexOCR-Train,我们使用监督微调(SFT)和基于可验证奖励(来自LaTeX单元测试)的强化学习训练了一个20亿参数模型TexOCR,直接强化编译性和引用完整性。在21个前沿模型上的实验表明,现有系统常违反关键文档不变量,包括一致的章节结构、正确的浮动元素位置和有效的标签-引用链接,严重影响编译可靠性与下游可用性。分析显示,采用可验证奖励的强化学习相比SFT显著提升结构与编译性能。
原文摘要 · Abstract (English)
Existing document OCR largely targets plain text or Markdown, discarding the structural and executable properties that make LaTeX essential for scientific publishing. We study page-level reconstruction of scientific PDFs into compilable LaTeX and introduce TexOCR-Bench, a benchmark, and TexOCR-Train, a large-scale training corpus, for this task. TexOCR-Bench features a multi-dimensional evaluation suite that jointly assesses transcription fidelity, structural faithfulness, and end-to-end compilability. Leveraging TexOCR-Train, we train a 2B-parameter model, TexOCR, using supervised fine-tuning (SFT) and reinforcement learning (RL) with verifiable rewards derived from LaTeX unit tests that directly enforce compilability and referential integrity. Experiments across 21 frontier models on TexOCR-Bench show that existing systems frequently violate key document invariants, including consistent section structure, correct float placement, and valid label-reference links, which undermines compilation reliability and downstream usability. Our analysis further reveals that RL with verifiable rewards yields consistent improvements over SFT alone, particularly on structural and compilation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。