arXiv:2602.03370cs.CVcs.LG2026-02

用扩散模型逐步优化手写公式符号与结构,提升识别准确率。

GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition

  • 将公式识别转为符号迭代修正,避免逐个生成的偏差
  • 在MathWriting上达5.51%错误率,59.9%表达式正确率
  • 对不同书写风格和跨数据集均表现稳健,适合实际应用

手写数学公式识别需处理多样符号与结构关系,但自回归模型存在暴露偏差与语法不一致问题。本文提出GryphOne,一种离散扩散框架,将HMER重构为符号迭代精炼过程,而非序列生成。该方法逐步优化符号及其关系,消除自回归依赖,提升识别一致性。通过符号感知分词与随机掩码互学习机制,增强对书写差异的鲁棒性。在MathWriting基准上,GryphOne实现5.51%的字符错误率(CER)与59.9%的表达式正确率(ExpRate),优于所有重实现模型及商用系统。在CROHME 2014-2023的独立测试中,亦展现强跨数据集泛化能力。

原文摘要 · Abstract (English)

Handwritten mathematical expression recognition (HMER) requires reasoning over diverse symbols and structures, yet autoregressive models struggle with exposure bias and syntax inconsistency. We present GryphOne, a discrete diffusion framework which reformulates HMER as iterative symbolic refinement instead of sequential generation. GryphOne progressively refines symbols and relations, removing autoregression and improving consistency. Symbol-aware tokenization and random-masking mutual learning further enhance robustness to handwriting diversity. On the MathWriting benchmark, GryphOne achieves 5.51% CER and 59.9% EM (ExpRate), outperforming all reimplemented models in the matched setting as well as the commercial HMER system. Held-out evaluation on CROHME 2014-2023 further shows strong cross-dataset generalization.

公式识别扩散模型符号优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。