arXiv:2604.17680cs.IR2026-04

构建首个AI/ML领域必引文献推荐基准,解决关键文献遗漏问题

MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature

  • 基于15个顶会15万篇论文构建候选池,标注三类必引特征
  • 提出三层次标注体系,用大模型辅助+人工验证确保质量
  • 评测需从标题摘要中召回必引论文,现有方法仍不理想

人工智能与机器学习文献爆炸式增长,顶会如NeurIPS和ICLR每年接收数千篇论文,研究者全面引用变得愈发困难。尽管引文推荐研究已有十余年,但现有系统多关注广泛相关性,而非识别真正‘必引’文献——即实验基线、基础方法和核心依赖。本文提出MasterSet,一个专为AI/ML领域设计的大规模必引推荐基准。该数据集涵盖15个顶级会议的超15万篇论文,作为检索候选池。通过三层次标注体系:(I) 实验基线状态,(II) 核心相关性(1–5分),(III) 论文中提及频率。标注流程采用大模型判断,并经人工专家在分层样本上验证。任务要求仅根据查询论文的标题与摘要,从候选池中检索必引文献,评估指标为Recall@K。我们建立稀疏检索、密集科学嵌入与图方法基线,证明必引文献检索仍是开放难题。

原文摘要 · Abstract (English)

The explosive growth of AI and machine learning literature -- with venues like NeurIPS and ICLR now accepting thousands of papers annually -- has made comprehensive citation coverage increasingly difficult for researchers. While citation recommendation has been studied for over a decade, existing systems primarily focus on broad relevance rather than identifying the critical set of ``must-cite'' papers: direct experimental baselines, foundational methods, and core dependencies whose omission would misrepresent a contribution's novelty or undermine reproducibility. We introduce MasterSet, a large-scale benchmark specifically designed to evaluate must-cite recommendation in the AI/ML domain. MasterSet incorporates over 150,000 papers collected from official conference proceedings/websites of 15 leading venues, serving as a comprehensive candidate pool for retrieval. We annotate citations with a three-tier labeling scheme: (I) experimental baseline status, (II) core relevance (1--5 scale), and (III) intra-paper mention frequency. Our annotation pipeline leverages an LLM-based judge, validated by human experts on a stratified sample. The benchmark task requires retrieving must-cite papers from the candidate pool given only a query paper's title and abstract, evaluated by Recall@$K$. We establish baselines using sparse retrieval, dense scientific embeddings, and graph-based methods, demonstrating that must-cite retrieval remains a challenging open problem.

引文推荐基准测试AI文献必引文献

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。