arXiv:2606.20235cs.IRcs.AI2026-06

构建开放环境下的学术论文搜索评估基准,提升智能搜索代理的评测标准。

ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments

论文配图:ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments
图 1 · 摘自论文原文
  • 基于1000+计算机科学主题与4类研究意图构建评估体系。
  • 最佳代理仅达Recall@100 0.314,显示仍有巨大提升空间。
  • 支持多维度评估,适合研究智能检索系统者使用。

学术论文搜索是科研核心环节,基于大模型的搜索代理正成为一种有前景的迭代式、目标驱动的文献探索范式。然而,现有基准难以在真实开放文献环境下系统评估智能体的学术搜索能力。我们提出ScholarQuest,一个大规模、分类引导的智能学术论文搜索评估基准。该基准涵盖超过1,000个计算机科学主题和四类典型研究意图:方法导向、场景锚定、比较型与范围控制型查询。同时提供可扩展的答案构造方式与共享检索后端ScholarBase,支持可复现评估。基准测试结果显示,智能体方法优于单次检索基线,但表现有限——最佳代理仅达Recall@100 0.314、Recall@All 0.355,表明仍有显著改进空间。对搜索效率、意图鲁棒性及失败案例的分析进一步验证了该基准提供多维评估信号的能力。

原文摘要 · Abstract (English)

Academic paper search is a core step in scientific research, and LLM-based search agents are emerging as a promising paradigm for iterative, intent-driven literature exploration. However, existing benchmarks are insufficient for systematically evaluating agentic academic search under realistic open literature environments. We propose ScholarQuest, a large-scale, taxonomy-guided benchmark for agentic academic paper search. ScholarQuest is constructed from over 1,000 computer science topics and four representative research intents, including method-oriented, setting-anchored, comparison-based, and scope-controlled queries. It further provides scalable answer construction and a shared retrieval backend ScholarBase for reproducible evaluation. Benchmarking results show that agentic methods outperform single-shot retrieval baselines, yet the best-performing agent only achieves 0.314 Recall@100 and 0.355 Recall@All, indicating substantial room for improvement. In addition, analyses of search efficiency, intent-level robustness, and failure cases further highlight the benchmark's ability to provide multi-dimensional evaluation signals for academic paper search agents.

学术搜索智能代理评估基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。