构建学术研究信息搜集评估框架,揭示大模型在深度研究中的关键瓶颈
ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
- 分阶段评估大模型的文献检索能力,拆解查询规划、工具调用与相关性判断
- 迭代查询分解使F1提升2.9至3.3倍,思维扩展牺牲召回率换取精确度
- 查询规划与相关性判断是决定闭源与开源模型差距的核心瓶颈
大语言模型已从单轮问答演进为可迭代分解研究问题、调用检索工具并跨轮次整合信息的深度研究系统。现有评估多采用整体报告评分,但此端到端范式将模型决策、流程设计与环境反馈紧密耦合,难以分解分析各组件表现。我们提出ScholarGym,一个隔离深度研究中信息搜集阶段的评估环境。在统一工作流下,ScholarGym将研究过程显式划分为查询规划、工具调用与相关性评估三阶段,并在包含57万篇论文的静态语料库上,基于确定性检索对2,536个专家标注查询进行评估。系统实验表明:迭代查询分解相较单查询检索可带来2.9至3.3倍的F1提升;具有长思考能力的模型以牺牲召回率为代价换取更高精度;查询规划质量与相关性评估共同构成制约闭源与开源模型性能差异的双重瓶颈。
原文摘要 · Abstract (English)
Large language models have advanced from single-turn question answering to deep research systems that iteratively decompose research questions, invoke retrieval tools, and synthesize information across multiple rounds. Evaluating such systems typically involves scoring their final research reports holistically, but this end-to-end paradigm tightly couples the language model's decision-making, workflow design, and environmental feedback, precluding decomposable analysis of individual components. We introduce ScholarGym, an evaluation environment that isolates the information-gathering stage of deep research on academic literature. Under a unified workflow, ScholarGym decomposes the research process into three explicit stages -- Query Planning, Tool Invocation, and Relevance Assessment -- and evaluates each against 2,536 expert-annotated queries over a static corpus of 570K papers with deterministic retrieval. Systematic experiments reveal that iterative query decomposition yields 2.9--3.3$\times$ F1 gains over single-query retrieval, models with extended thinking trade recall for precision, and Query Planning quality together with Relevance Assessment constitute dual bottlenecks that separate proprietary from open-source model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。