构建多角度论文检索基准,测试模型能否从不同视角找到同一论文。
Can Retrievers Find the Same Paper from Different Aspects? A Multi-Aspect Full-Paper Scientific Retrieval Benchmark

- 设计多角度检索评测框架,覆盖动机、方法、实验等论文维度。
- 最强模型仅15.7%能从所有角度正确召回论文,表明现有系统仍不完整。
- 适合研究科学信息检索、论文推荐与AI助研系统的学者使用。
科学论文包含背景、方法、实验结果等多个可检索维度,但现有检索评测多仅关注单个查询-论文相关性,忽视了同一论文在不同视角下的表现。为此,我们提出MAPLE——一个专家验证的多角度全论文检索基准,评估检索器能否从动机、方法、实验发现等不同方面一致地召回同一论文。MAPLE包含2095个关于近期机器学习与自然语言处理论文的查询,基于文本与多模态内容构建。我们进一步提出MAPLE-Synth,一种基于OpenReview讨论和人工撰写示例的检索式上下文学习管道,生成反映研究人员真实兴趣的自然查询。专家评估表明,这些生成查询与人工撰写查询在真实性和相关性上相当。跨词汇、科学领域、通用文本及多模态检索器的实验显示:从任一角度检索成功(AnyAspect@20达98.1%),但同时从所有角度检索成功率仅为15.7%。结果查询和表格引用查询尤为困难。尽管多块聚合有所提升,仍存在大量失败案例。MAPLE为评估与开发更全面表征科学论文的检索器提供了测试平台。
原文摘要 · Abstract (English)
Scientific papers contain multiple searchable facets such as background, methods. However, many paper retrieval benchmarks merely evaluate individual query-paper relevance, while overlooking other facets of the same paper. To bridge this gap, we introduce MAPLE, an expert-validated benchmark for multi-aspect, full-paper retrieval that evaluates whether retrievers can consistently recover the same paper from queries targeting its motivation, method, and experimental findings. MAPLE contains 2,095 queries about recent ML and NLP papers, grounded in both textual and multimodal content. We further propose MAPLE-Synth, a retrieval-based in-context learning pipeline that leverages OpenReview discussions and human-written query exemplars to generate realistic queries reflecting researchers' interests in different aspects of a paper. Our expert validation shows that these queries are comparable in realism to human-written queries and highly relevant to the target papers. Experiments across lexical, scientific-domain, general-purpose text, and multimodal retrievers reveal a substantial gap between retrieving a paper from any one aspect and retrieving it from all aspects: the strongest model achieves 98.1% AnyAspect@20 but only 15.7% AllAspect@20. Experiment/result queries and table-referenced queries are particularly difficult across retrievers. Although multi-chunk aggregation improves multi-aspect paper retrieval, considerable failures persist. MAPLE provides a testbed for evaluating and developing retrievers that represent scientific papers more comprehensively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。