arXiv:2505.19253cs.IR2025-05被引 51

开源可复现的深度研究评估平台,替代昂贵商业搜索接口。

DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research

  • 用公开网络数据构建可复现搜索接口,支持高效检索
  • 在多个指标上表现与商业接口相当,排名一致
  • 适合研究者低成本训练和验证智能搜索系统

深度研究系统是一类新兴的代理式信息检索方法,能针对复杂问题生成全面且有依据的报告。然而,现有框架大多依赖动态的商业搜索API,导致可复现性差、透明度低且成本高昂。为此,我们提出 extsc{DeepResearchGym}——一个开源沙箱,结合可复现的搜索API与严格的评估协议,用于基准测试深度研究系统。该API基于ClueWeb22和FineWeb两大公开语料库,使用先进的密集检索器与DiskANN近似最近邻搜索,实现比主流商业API更低延迟,并保证跨运行文档排序稳定,且对科研免费。为评估输出质量,我们在Researchy Questions基准上引入大模型作为裁判(LLM-as-a-judge),自动衡量用户需求对齐度、检索忠实度与报告质量。实验表明,集成 extsc{DeepResearchGym}的系统性能与使用商业API相当,不同评估指标下性能排名保持一致。对短答案搜索代理的案例研究进一步证明,该沙箱可有效支持低成本训练,且训练模型具备迁移到商业搜索的能力。

原文摘要 · Abstract (English)

Deep research systems represent an emerging class of agentic information retrieval methods that generate comprehensive and well-supported reports to complex queries. However, most existing frameworks rely on dynamic commercial search APIs, which pose reproducibility and transparency challenges in addition to their cost. To address these limitations, we introduce \textsc{DeepResearchGym} as an open-source sandbox that combines a reproducible search API with a rigorous evaluation protocol for benchmarking deep research systems. The API indexes large-scale public web corpora, namely ClueWeb22 and FineWeb, using a state-of-the-art dense retriever and approximate nearest neighbor search via DiskANN. It achieves lower latency than popular commercial APIs while ensuring stable document rankings across runs, and is free for research use. To evaluate deep research systems' outputs, we extend the Researchy Questions benchmark with automatic metrics through LLM-as-a-judge to measure alignment with users' information needs, retrieval faithfulness, and report quality. Experimental results show that systems integrated with~\textsc{DeepResearchGym} achieve performance comparable to those using commercial APIs, with performance rankings remaining consistent across evaluation metrics. A case study on short-answer search agents further demonstrates the sandbox's utility for cost-effective training, showing that models trained within the sandbox can generalize to commercial search.

信息检索可复现性开源工具评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。