arXiv:2512.13731cs.CVcs.AI2025-12AAAI被引 2

针对复杂数学表达式识别难题,构建了新基准与数据集,提出高效模型。

Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

  • 设计三难度分级基准CMER-Bench,评估现有模型表现。
  • 在300万样本数据集上训练的模型,复杂表达式识别准确率显著提升。
  • 适合需要高精度数学表达识别的研究者与开发者使用。

数学表达式识别(MER)在简单表达式上已取得显著进展,但对包含多个符号和多行结构的复杂表达式仍面临巨大挑战。本文首次提出CMER-Bench,一个精心构建的基准,将表达式分为简单、中等、复杂三个难度等级。基于该基准,我们全面评估了现有MER模型及通用多模态大模型(MLLMs)。结果表明,当前方法在简单与中等表达式上表现良好,但在复杂表达式上性能大幅下降,主要因现有公开训练数据以简单样本为主。为此,我们构建了大规模数据集MER-17M和CMER-3M,专注于复杂表达式识别。这些数据集提供丰富多样的样本,支持高精度、鲁棒的复杂MER模型发展。此外,为应对复杂表达式的复杂空间布局,我们提出一种新型表达式分词器,以及一种名为“结构化数学语言”的新表示形式,显式建模表达式的层级与空间结构,超越传统LaTeX格式。基于此,我们提出了专用于复杂表达式识别的模型CMERNet,采用编码器-解码器架构,并在CMER-3M上训练。实验表明,仅含12500万参数的CMERNet,在CMER-Bench上显著优于现有MER模型和MLLMs。

原文摘要 · Abstract (English)

Mathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many tokens and multiple lines remains a formidable challenge. In this paper, we first introduce CMER-Bench, a carefully constructed benchmark that categorizes expressions into three difficulty levels: easy, moderate, and complex. Leveraging CMER-Bench, we conduct a comprehensive evaluation of existing MER models and general-purpose multimodal large language models (MLLMs). The results reveal that while current methods perform well on easy and moderate expressions, their performance degrades significantly when handling complex mathematical expressions, mainly because existing public training datasets are primarily composed of simple samples. In response, we propose MER-17M and CMER-3M that are large-scale datasets emphasizing the recognition of complex mathematical expressions. The datasets provide rich and diverse samples to support the development of accurate and robust complex MER models. Furthermore, to address the challenges posed by the complicated spatial layout of complex expressions, we introduce a novel expression tokenizer, and a new representation called Structured Mathematical Language, which explicitly models the hierarchical and spatial structure of expressions beyond LaTeX format. Based on these, we propose a specialized model named CMERNet, built upon an encoder-decoder architecture and trained on CMER-3M. Experimental results show that CMERNet, with only 125 million parameters, significantly outperforms existing MER models and MLLMs on CMER-Bench.

数学表达识别数据集构建模型优化结构化表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。