arXiv:2502.05557cs.CV2025-02被引 1

融合CNN与Transformer优势,提升手写公式识别准确率

MMHMER:Multi-viewer and Multi-task for Handwritten Mathematical Expression Recognition

  • 设计多视角多任务框架,协同使用CNN与Transformer
  • 在CROHME14/16/19数据集上分别达63.96%/62.51%/65.46%准确率
  • 性能优于Posformer,适合需要高精度的数学表达式识别场景

手写数学表达式识别(HMER)已取得显著进展,现有方法多基于CNN/RNN结合GRU或Transformer架构,各有优劣。本文提出一种高效的CNN-Transformer多视角、多任务框架,以融合两者优势。通过利用CNN的特征提取能力与Transformer的序列建模能力,模型能更好处理手写公式的复杂性。在CROHME14、CROHME16和CROHME19数据集上,该模型分别达到63.96%、62.51%和65.46%的表达式识别率(ExpRate),相较于Posformer分别提升1.28%、1.48%和0.58%。

原文摘要 · Abstract (English)

Handwritten Mathematical Expression Recognition (HMER) methods have made remarkable progress, with most existing HMER approaches based on either a hybrid CNN/RNN-based with GRU architecture or Transformer architectures. Each of these has its strengths and weaknesses. Leveraging different model structures as viewers and effectively integrating their diverse capabilities presents an intriguing avenue for exploration. This involves addressing two key challenges: 1) How to fuse these two methods effectively, and 2) How to achieve higher performance under an appropriate level of complexity. This paper proposes an efficient CNN-Transformer multi-viewer, multi-task approach to enhance the model's recognition performance. Our MMHMER model achieves 63.96%, 62.51%, and 65.46% ExpRate on CROHME14, CROHME16, and CROHME19, outperforming Posformer with an absolute gain of 1.28%, 1.48%, and 0.58%. The main contribution of our approach is that we propose a new multi-view, multi-task framework that can effectively integrate the strengths of CNN and Transformer. By leveraging the feature extraction capabilities of CNN and the sequence modeling capabilities of Transformer, our model can better handle the complexity of handwritten mathematical expressions.

手写识别多模态融合数学表达式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。