arXiv:2601.08834cs.CVcs.AI2026-01被引 7

通过解耦格式的强化学习,提升复杂文档OCR的准确率

Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR

  • 基于熵值筛选格式密集型文本,针对性优化
  • 在OmniDocBench上达90.41分,创端到端模型新纪录
  • 适合研究文档理解、OCR鲁棒性与强化学习融合者

通过OCR从图像或扫描文档中读取文本一直是研究重点。直观上,文本阅读被视为简单的感知任务,现有工作主要聚焦于构建丰富数据以增强监督微调能力。本文观察到,即使先进OCR模型在格式化文本(如公式、表格等)上的输出熵值也显著高于普通文本,差距可达一个数量级。这一统计特征表明,先进OCR在处理格式敏感文档时存在高不确定性,提示需通过多路径推理来提升性能。为此,我们提出格式解耦强化学习(FD-RL),利用高熵模式进行定向优化。方法采用基于熵的数据过滤策略识别格式密集实例,并设计针对不同格式类型的解耦奖励机制,实现格式层面验证而非逐标记记忆。FD-RL在OmniDocBench上取得平均90.41分,创下端到端模型新纪录。此外,我们对数据、训练、过滤与奖励策略进行了全面消融实验,充分验证其有效性。

原文摘要 · Abstract (English)

Reading text from images or scanned documents via OCR models has been a longstanding focus of researchers. Intuitively, text reading is perceived as a straightforward perceptual task, and existing work primarily focuses on constructing enriched data engineering to enhance SFT capabilities. In this work, we observe that even advanced OCR models exhibit significantly higher entropy in formatted text (\emph{e.g.}, formula, table, etc.) compared to plain text, often by an order of magnitude. These statistical patterns reveal that advanced OCR models struggle with high output uncertainty when dealing with format sensitive document, suggesting that reasoning over diverse reading pathways may improve OCR performance. To address this, we propose format decoupled reinforcement learning (FD-RL), which leverages high-entropy patterns for targeted optimization. Our approach employs entropy-based data filtration strategy to identify format-intensive instances, and adopt format decoupled rewards tailored to different format types, enabling format-level validation rather than token-level memorization. FD-RL achieves an average score of 90.41 on OmniDocBench, setting a new record for end-to-end models on this highly popular benchmark. More importantly, we conduct comprehensive ablation studies over data, training, filtering, and rewarding strategies, thoroughly validating their effectiveness.

OCR强化学习格式理解文档分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。