arXiv:2605.29234cs.AIcs.IR2026-05被引 1

改进文献检索流程,发现人工参考文献非可靠基准。

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

论文配图:Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth
图 1 · 摘自论文原文
  • 构建深度研究流水线,基于论文引文广度扩展检索结果
  • 召回率从不足20%提升至80%以上,显著优于纯API搜索
  • 人工引用仅51%被判定为相关,适合评估多维度指标协同

我们从两个互补角度研究大规模文献检索:优化检索流程,以及检验人工参考文献列表作为评估基准的可靠性。首先,我们实现了一种深度研究流水线,通过分析查询论文并沿其参考文献进行广度优先扩展,显著优于单纯的API搜索,在RollingEval-Jun25(一个包含250篇论文的文献检索基准)上将召回率从低于20%提升至超过80%。其次,我们使用中立的LLM作为评判者,评估人类引用是否适合作为任务的黄金标准。结果发现:仅有51%的人工引用被判定为中等及以上相关性,而最强的AI重排序模型达到86%–88%。在OpenAlex合作者图谱上的分析表明,人类引用直接合作者的概率是最佳AI重排序器的2.5倍。综合结果表明,单一维度的文献检索评估不可靠,应联合报告召回率、主题相关性评分、排序列表多样性及合作者距离诊断等多维指标。

原文摘要 · Abstract (English)

We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as an evaluation target. First, we implement a Deep Research pipeline that processes the full query paper and expands the retrieved results breadth-first along their bibliographies, and show that it substantially outperforms vanilla API-only search, raising recall on RollingEval-Jun25 (a 250-paper literature-search benchmark) from below 20% to above 80%. Second, we use a neutral LLM-as-a-judge to determine if human references are sound ground truth for the task. We find significant limitations: only 51% of human citations are judged moderately relevant or higher, against 86--88% for the strongest AI-based re-rankers. We study this gap on the OpenAlex co-authorship graph, finding that humans are 2.5x more likely than the best AI re-rankers to cite a direct collaborator. Together, our results argue against single-axis literature-search evaluation: recall, topical-relevance scoring, ranked-list diversity, and a co-authorship-distance diagnostic each measure complementary properties of citation quality and should be reported jointly.

文献检索AI评估引用质量多维评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。