新基准让深度研究模型评估更公平透明,可精准分析检索与推理能力。
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent

- 基于固定语料库和人工验证文档,实现可复现的评测
- 开源模型仅3.86%准确率,GPT-5达55.9%,融合检索器后提升至70.1%
- 适合研究检索系统、推理能力或可解释性的人群使用
深度研究代理(Deep-Research agents)结合大语言模型(LLMs)与搜索工具,在处理需迭代搜索规划和结果推理的复杂查询上表现优异。现有基准如BrowseComp依赖黑盒实时网络搜索接口,存在公平性不足(动态接口阻碍公平比较与复现)和透明度低(无法控制文档语料,难以分离检索器贡献)的问题。为此,我们提出BrowseComp-Plus,基于BrowseComp构建,采用固定且精心筛选的语料库,每个查询配有经人工验证的支持文档及挖掘出的挑战性负样本,支持受控实验。该基准能有效区分不同深度研究系统性能:例如,开源模型Search-R1搭配BM25检索器准确率为3.86%,而GPT-5达到55.9%;进一步结合Qwen3-Embedding-8B检索器,准确率提升至70.1%,且减少搜索调用次数。该基准支持对深度研究代理与检索方法的全面评估与解耦分析,助力理解检索有效性、引用准确性与上下文工程在深度研究系统中的作用。
原文摘要 · Abstract (English)
Deep-Research agents, which integrate large language models (LLMs) with search tools, have shown success in improving the effectiveness of handling complex queries that require iterative search planning and reasoning over search results. Evaluations on current benchmarks like BrowseComp relies on black-box live web search APIs, have notable limitations in (1) fairness: dynamic and opaque web APIs hinder fair comparisons and reproducibility of deep research methods; (2) transparency: lack of control over the document corpus makes it difficult to isolate retriever contributions. In other words, the current evaluations may compare a complete deep research system at a given time, but they do not foster well-controlled experiments to provide insights into the capability of underlying deep research LLMs. To address these challenges, we introduce BrowseComp-Plus, a benchmark derived from BrowseComp, employing a fixed, carefully curated corpus. Each query in BrowseComp-Plus includes human-verified supporting documents and mined challenging negatives, enabling controlled experimentation. The benchmark is shown to be effective in distinguishing the performance of deep research systems. For instance, the open-source model Search-R1, when paired with the BM25 retriever, achieves 3.86% accuracy, whereas the GPT-5 achieves 55.9%. Integrating the GPT-5 with the Qwen3-Embedding-8B retriever further enhances its accuracy to 70.1% with fewer search calls. This benchmark allows comprehensive evaluation and disentangled analysis of deep research agents and retrieval methods, fostering insights into retrieval effectiveness, citation accuracy, and context engineering in Deep-Research system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。