首个专用于评估大模型学术搜索能力的数据集,挑战真实科研场景下的信息检索。
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
- 构建覆盖15个学科的学术搜索数据集,贴近真实研究需求
- 单模型难以直接回答,需至少三次深度搜索才能得出答案
- 强调结果可验证性,提供清晰来源与简明解析,便于审计
大型语言模型(LLMs)的信息检索能力受到广泛关注。现有基准如OpenAI的BrowseComp主要针对通用搜索场景,未能充分应对学术搜索的特殊需求,包括深层次文献追溯与组织、专业学术数据库支持、长尾学术知识导航以及学术严谨性保障。为此,我们提出ScholarSearch,这是首个专门用于评估大模型在学术研究中复杂信息检索能力的数据集。该数据集具备四大特性:学术实用性,问题内容紧密贴合真实学术学习与研究环境,避免刻意误导模型;高难度性,答案对单一模型(如Grok DeepSearch或Gemini Deep Research)而言极具挑战,通常需要至少三次深度搜索才能推导得出;简洁评估性,通过限定条件使答案尽可能唯一,附带明确来源与简要解题说明,极大提升后续审计与验证效率,优于当前国内外缺乏分析的搜索数据集;广泛覆盖性,涵盖至少15个不同学术领域。通过ScholarSearch,我们期望更精准地衡量并推动大模型在复杂学术信息检索任务中的性能提升。数据集已公开:https://huggingface.co/datasets/PKU-DS-LAB/ScholarSearch
原文摘要 · Abstract (English)
Large Language Models (LLMs)' search capabilities have garnered significant attention. Existing benchmarks, such as OpenAI's BrowseComp, primarily focus on general search scenarios and fail to adequately address the specific demands of academic search. These demands include deeper literature tracing and organization, professional support for academic databases, the ability to navigate long-tail academic knowledge, and ensuring academic rigor. Here, we proposed ScholarSearch, the first dataset specifically designed to evaluate the complex information retrieval capabilities of Large Language Models (LLMs) in academic research. ScholarSearch possesses the following key characteristics: Academic Practicality, where question content closely mirrors real academic learning and research environments, avoiding deliberately misleading models; High Difficulty, with answers that are challenging for single models (e.g., Grok DeepSearch or Gemini Deep Research) to provide directly, often requiring at least three deep searches to derive; Concise Evaluation, where limiting conditions ensure answers are as unique as possible, accompanied by clear sources and brief solution explanations, greatly facilitating subsequent audit and verification, surpassing the current lack of analyzed search datasets both domestically and internationally; and Broad Coverage, as the dataset spans at least 15 different academic disciplines. Through ScholarSearch, we expect to more precisely measure and promote the performance improvement of LLMs in complex academic information retrieval tasks. The data is available at: https://huggingface.co/datasets/PKU-DS-LAB/ScholarSearch
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。