首个数学信息检索统一基准,评估模型在数学文档中的多项检索能力。
MIRB: Mathematical Information Retrieval Benchmark
- 构建涵盖4类任务的12个数据集,覆盖数学语义、问答、前提与公式检索
- 评估13种模型表现,揭示数学检索任务的核心挑战
- 适合研究数学AI、自动定理证明和知识库系统的研究者参考
数学信息检索(MIR)是从数学文档中检索信息的任务,在数学库定理搜索、数学论坛问答和自动化定理证明中的前提选择等应用中起关键作用。然而,缺乏统一的评估基准来衡量这些多样化的检索任务。本文提出MIRB(数学信息检索基准),用于评估检索模型在MIR任务上的能力。MIRB包含四个任务:语义命题检索、问答检索、前提检索和公式检索,共涵盖12个数据集。我们在该基准上评估了13种检索模型,并分析了MIR固有的挑战。我们希望MIRB能为MIR系统提供一个全面的评估框架,推动更适用于数学领域的高效检索模型的发展。
原文摘要 · Abstract (English)
Mathematical Information Retrieval (MIR) is the task of retrieving information from mathematical documents and plays a key role in various applications, including theorem search in mathematical libraries, answer retrieval on math forums, and premise selection in automated theorem proving. However, a unified benchmark for evaluating these diverse retrieval tasks has been lacking. In this paper, we introduce MIRB (Mathematical Information Retrieval Benchmark) to assess the MIR capabilities of retrieval models. MIRB includes four tasks: semantic statement retrieval, question-answer retrieval, premise retrieval, and formula retrieval, spanning a total of 12 datasets. We evaluate 13 retrieval models on this benchmark and analyze the challenges inherent to MIR. We hope that MIRB provides a comprehensive framework for evaluating MIR systems and helps advance the development of more effective retrieval models tailored to the mathematical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。