arXiv:2606.23959cs.CLcs.LG2026-06

测试嵌入模型能否识别数学等价命题,发现现有模型依赖术语而非本质。

Does My Embedding Reflect That $A = B$? Evaluating Mathematical Equivalence in Embedding Models

论文配图:Does My Embedding Reflect That $A = B$? Evaluating Mathematical Equivalence in Embedding Models
图 1 · 摘自论文原文
  • 构建MELD数据集,包含语义等价但表达差异大的数学命题对
  • 现有模型多按术语分组,无法捕捉数学本质等价性
  • 提出对比学习方法,提升跨形式化表达的检索能力

由于数学高度抽象,同一命题在不同领域中可能呈现迥异表述。历史上许多突破源于发现不同领域的命题实际等价。随着形式化资源增多,高效可靠地在数学‘语言’间导航(如从Lean到自然语言)的需求日益迫切。本文探究当前嵌入模型是否能捕捉数学等价性。为此,我们构建了数学等价但词汇不同的命题对数据集MELD。实验表明,现有先进嵌入模型倾向于根据表述术语分组,而非底层数学内容。基于此,我们提出一种对比学习方法,聚焦于对齐非形式化陈述与不同形式化表达。实验显示,该方法不仅提升了非形式化-形式化检索任务表现,还在仅含自然语言的MELD数据集上取得显著改进。

原文摘要 · Abstract (English)

Because mathematics is highly abstract, a single statement can take very different forms depending on what subfield it is framed in. There are many examples where breakthroughs occurred after researchers discovered that a question had already been answered in a different field. At the same time, the growth of new resources related to formalization has increased the need for tools that enable efficient and reliable navigation between mathematical 'languages' (e.g., from Lean to natural language). In this paper, we investigate whether current embedding models capture mathematical equivalence. To do this, we introduce the Mathematically Equivalent but Lexically Different Pairs (MELD) Dataset, a collection of mathematically equivalent statements that are expressed in very different language. We show that current state-of-the-art embedding models tend to group statements by the terminology used to make them instead of the underlying math. Motivated by this, we propose a contrastive approach to learning embeddings of mathematical text that focuses on aligning informal statements with different formalizations. Our experiments demonstrate that this leads to improvements not only on informal-formal retrieval tasks but also on MELD, which only contains natural language statements.

数学推理嵌入模型语义等价对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。