arXiv:2508.00311cs.CV2025-08被引 2

用通用视觉语言模型实现复杂公式识别,准确率超越专用模型

DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios

  • 基于通用视觉语言模型构建统一框架,无需专用架构
  • 在多学科复杂布局中达到当前最优识别精度
  • 适合科研文献智能分析与自动化处理场景

数学公式的光学字符识别对科学文献的智能分析至关重要。然而,任务特定及通用视觉-语言模型常难以应对数学内容固有的结构多样性、复杂性与现实世界中的变化。本文提出DocTron-Formula,一个基于通用视觉-语言模型的统一框架,无需专门设计架构。同时引入CSFormula数据集,涵盖跨学科、多层次(行、段、页)结构复杂的公式。通过简单的监督微调,该方法在多种风格、科学领域和复杂版式下均实现领先性能。实验表明,本方法不仅在准确率与鲁棒性上超越专用模型,更建立了一种复杂科学文档自动化理解的新范式。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) for mathematical formula is essential for the intelligent analysis of scientific literature. However, both task-specific and general vision-language models often struggle to handle the structural diversity, complexity, and real-world variability inherent in mathematical content. In this work, we present DocTron-Formula, a unified framework built upon general vision-language models, thereby eliminating the need for specialized architectures. Furthermore, we introduce CSFormula, a large-scale and challenging dataset that encompasses multidisciplinary and structurally complex formulas at the line, paragraph, and page levels. Through straightforward supervised fine-tuning, our approach achieves state-of-the-art performance across a variety of styles, scientific domains, and complex layouts. Experimental results demonstrate that our method not only surpasses specialized models in terms of accuracy and robustness, but also establishes a new paradigm for the automated understanding of complex scientific documents.

公式识别视觉语言模型OCR科学文献

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。