arXiv:2602.05216cs.IRcs.AI2026-02被引 8

在900万数学定理上实现精准语义搜索,助力科研高效查找。

Semantic Search over 9 Million Mathematical Theorems

论文配图:Semantic Search over 9 Million Mathematical Theorems
图 1 · 摘自论文原文
  • 用自然语言描述定理,构建可检索的语义表示
  • 在专业数学家查询集上,定理与论文检索效果显著优于基线
  • 适合数学研究者、AI推理系统快速定位核心定理

数学结果的搜索仍面临挑战:现有工具多返回整篇论文,而数学家和定理证明代理常需精确查找特定定理、引理或命题。尽管语义搜索进展迅速,但其在大型高技术性语料(如研究级数学定理)上的表现仍不明确。本文首次在涵盖920万条定理陈述的统一语料库上系统研究语义定理检索,该语料来自arXiv及其他七项来源,是目前公开的最大规模人类撰写的研究级定理论文集。我们为每条定理构建短自然语言描述作为检索表示,并系统分析表示上下文、语言模型选择、嵌入模型及提示策略对检索质量的影响。在由专业数学家撰写的精选定理搜索查询集上,本方法在定理级与论文级检索上均显著优于现有基线,证明了大规模语义定理搜索在互联网规模下的可行性和有效性。项目主页、搜索工具、数据集、REST API及MCP服务器已开放于theoremsearch.com。

原文摘要 · Abstract (English)

Searching for mathematical results remains difficult: most existing tools retrieve entire papers, while mathematicians and theorem-proving agents often seek a specific theorem, lemma, or proposition that answers a query. While semantic search has seen rapid progress, its behavior on large, highly technical corpora such as research-level mathematical theorems remains poorly understood. In this work, we introduce and study semantic theorem retrieval at scale over a unified corpus of $9.2$ million theorem statements extracted from arXiv and seven other sources, representing the largest publicly available corpus of human-authored, research-level theorems. We represent each theorem with a short natural-language description as a retrieval representation and systematically analyze how representation context, language model choice, embedding model, and prompting strategy affect retrieval quality. On a curated evaluation set of theorem-search queries written by professional mathematicians, our approach substantially improves both theorem-level and paper-level retrieval compared to existing baselines, demonstrating that semantic theorem search is feasible and effective at web scale. The project page, search tool, dataset, REST API, and MCP server are available at theoremsearch.com.

数学定理语义搜索知识检索arXiv

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。