arXiv:2603.07929cs.CV2026-03中稿 · as oral presentati…被引 2

用混合视觉变压器提升数学表达式识别准确率

A Hybrid Vision Transformer Approach for Mathematical Expression Recognition

  • 采用带2D位置编码的混合视觉变压器提取符号间复杂关系
  • 在IM2LATEX-100K数据集上达89.94的BLEU分数,超越现有方法
  • 适合需要高精度数学公式识别的研究者与开发者

文档分析中的关键挑战之一是数学表达式识别。与仅关注一维结构的文本识别不同,数学表达式识别因具有二维结构和符号大小差异而更为复杂。本文提出使用带有2D位置编码的混合视觉变压器(HVT)作为编码器,从图像中提取符号间的复杂关系。采用覆盖注意力解码器以更好追踪注意力历史,缓解欠解析和过解析问题。还验证了利用ViT的[CLS]标记作为解码器初始嵌入的有效性。在IM2LATEX-100K数据集上的实验表明,该方法取得了89.94的BLEU分数,优于当前最优方法。

原文摘要 · Abstract (English)

One of the crucial challenges taken in document analysis is mathematical expression recognition. Unlike text recognition which only focuses on one-dimensional structure images, mathematical expression recognition is a much more complicated problem because of its two-dimensional structure and different symbol size. In this paper, we propose using a Hybrid Vision Transformer (HVT) with 2D positional encoding as the encoder to extract the complex relationship between symbols from the image. A coverage attention decoder is used to better track attention's history to handle the under-parsing and over-parsing problems. We also showed the benefit of using the [CLS] token of ViT as the initial embedding of the decoder. Experiments performed on the IM2LATEX-100K dataset have shown the effectiveness of our method by achieving a BLEU score of 89.94 and outperforming current state-of-the-art methods.

数学表达式识别视觉变压器序列生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。