arXiv:2509.19760cs.CV2025-09被引 14

用强化学习提升视觉语言模型的文档布局与阅读顺序分析能力

Logics-Parsing Technical Report

  • 引入强化学习优化文档布局和阅读顺序的端到端推理
  • 在1078张页面级PDF上达到当前最佳性能
  • 支持化学公式和手写中文等多类型数据,适合复杂文档处理

大型视觉语言模型(LVLM)的进展推动了文档解析任务的发展。相较于传统流水线方法,端到端范式通过集成光学字符识别(OCR)、表格识别、数学公式识别等能力,实现了从PDF图像到结构化输出的高效转换。然而,缺乏显式的文档布局与阅读顺序分析阶段,限制了LVLM在多栏报纸或海报等复杂文档类型上的表现。为此,本文提出Logics-Parsing:一种基于强化学习增强的端到端LVLM模型。该模型设计了精细的奖励机制,以优化复杂布局分析与阅读顺序推断。同时,通过在监督微调中引入化学公式和手写中文字符等多样化数据,提升了模型泛化能力。为实现严谨评估,我们构建了LogicsParsingBench,一个包含1,078张页面级PDF图像的数据集,覆盖九个主要类别及二十多个子类别,后续将公开。在该基准上的全面实验验证了所提模型在多种文档分析场景下的有效性与领先性能。

原文摘要 · Abstract (English)

Recent advances in Large Vision-Language models (LVLM) have spurred significant progress in document parsing task. Compared to traditional pipeline-based methods, end-to-end paradigms have shown their excellence in converting PDF images into structured outputs through integrated Optical Character Recognition (OCR), table recognition, mathematical formula recognition and so on. However, the absence of explicit analytical stages for document layouts and reading orders limits the LVLM's capability in handling complex document types such as multi-column newspapers or posters. To address this limitation, we propose in this report Logics-Parsing: an end-to-end LVLM-based model augmented with reinforcement learning. Our model incorporates meticulously designed reward mechanisms to optimize complex layout analysis and reading order inference. In addition, we expand the model's versatility by incorporating diverse data types such as chemical formulas and handwritten Chinese characters into supervised fine-tuning. Finally, to enable rigorous evaluation of our approach, we introduce LogicsParsingBench, a curated set of 1,078 page-level PDF images spanning nine major categories and over twenty sub-categories, which will be released later. Comprehensive experiments conducted on LogicsParsingBench have validated the efficacy and State-of-the-art (SOTA) performance of our proposed model across diverse document analysis scenarios. Project Page: https://github.com/alibaba/Logics-Parsing

文档解析强化学习视觉语言模型PDF处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。