arXiv:2608.02457cs.IRcs.AI2026-08

破解公式语法与语义的对应关系,提升科学文献检索效果

Syntax Meets Semantics: Understanding Scientific Formulae

论文配图:Syntax Meets Semantics: Understanding Scientific Formulae
图 1 · 摘自论文原文
  • 用图神经网络和文本编码器分别表征公式结构与语义
  • 发现原始表示中语法与语义对应微弱但潜在关联强
  • 通过对比学习对齐表示空间,显著提升跨模态检索

科学公式是学术交流的核心组成部分,其兼具结构化语法与语义承载双重特性,但在学术信息检索中仍缺乏系统研究。尽管已有研究表明联合建模语法与语义可提升检索性能,但二者底层表示间的关联尚未被深入探索。本文实证研究了公式语法与语义之间的跨模态对应关系。结果表明,尽管语法与语义的原始表示空间中可观测对应极弱,但存在较强的潜在相关性,揭示了两模态间存在显著表示错配。进一步评估了标准表示学习与对齐技术能否缓解此问题:采用基于图的编码器表征语法结构,文本编码器表征语义信息,并应用对比学习构建共享表示空间。实验显示,所学对齐显著提升了跨模态检索性能,说明显式表示学习可恢复原始表示中缺失的对应关系。

原文摘要 · Abstract (English)

Scientific formulae are a fundamental component of scholarly communication, yet their dual nature -- as structured syntax and carriers of semantics -- remains underexplored in scholarly information retrieval. Although prior studies show that jointly modeling syntactic and semantic modalities improves retrieval performance, the relationship between their underlying representations has not been systematically investigated. In this work, we empirically study cross-modal correspondence between formula syntax and semantics. We find that their native representation spaces exhibit extremely weak observable correspondence despite strong latent correlation, indicating a substantial representation mismatch between the two modalities. We further evaluate whether this mismatch can be reduced using standard representation learning and alignment techniques. We represent syntactic structure using graph-based encoders and semantic information using text-based encoders, then apply contrastive learning to induce a shared representation space. Results show that the learned alignment substantially improves cross-modal retrieval, suggesting that explicit representation learning can recover correspondence absent from the original representation spaces.

公式理解跨模态表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。