arXiv:2508.04162cs.IR2025-08被引 1

同时捕捉数学公式结构与语义,提升检索精度。

SSEmb: A Joint Structural and Semantic Embedding Framework for Mathematical Formula Retrieval

  • 用图对比学习编码公式结构,结合替换策略增强多样性。
  • 在ARQMath-3上,P'@10和nDCG'@10均超现有方法5个百分点以上。
  • 适合需要高精度数学公式检索的研究者或系统开发者。

数学信息检索中的公式检索是一项重要任务。我们提出SSEmb,一种能同时捕获数学公式结构与语义特征的嵌入框架。结构上,采用图对比学习对以算子图表示的公式进行编码,并通过替换策略实现图数据增强,以提升结构多样性并保持数学有效性。语义上,使用Sentence-BERT编码公式的上下文文本。对于每个查询及其候选公式,分别计算结构与语义相似度,并通过加权融合。在ARQMath-3公式检索任务中,SSEmb在P'@10和nDCG'@10上均优于现有嵌入方法超过5个百分点。此外,该方法可提升其他所有方法的表现,与Approach0结合后达到当前最优效果。

原文摘要 · Abstract (English)

Formula retrieval is an important topic in Mathematical Information Retrieval. We propose SSEmb, a novel embedding framework capable of capturing both structural and semantic features of mathematical formulas. Structurally, we employ Graph Contrastive Learning to encode formulas represented as Operator Graphs. To enhance structural diversity while preserving mathematical validity of these formula graphs, we introduce a novel graph data augmentation approach through a substitution strategy. Semantically, we utilize Sentence-BERT to encode the surrounding text of formulas. Finally, for each query and its candidates, structural and semantic similarities are calculated separately and then fused through a weighted scheme. In the ARQMath-3 formula retrieval task, SSEmb outperforms existing embedding-based methods by over 5 percentage points on P'@10 and nDCG'@10. Furthermore, SSEmb enhances the performance of all runs of other methods and achieves state-of-the-art results when combined with Approach0.

公式检索图神经网络语义嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。