构建首个深度研究代理评估基准,用专家标注的细粒度标准衡量推理与事实准确度。
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- 设计2500+条专家撰写的评分标准,匹配多样化真实问题
- 领先模型平均合规率不足68%,主因遗漏隐含上下文和推理不足
- 提供可复现的评估流程,适合研究智能研究助手的团队使用
深度研究(DR)是利用大语言模型应对开放性问题的新应用,需具备多步推理、跨文档整合及生成有证据支持的长篇回答能力。由于回答内容长且多样,存在多种有效解法并依赖动态信息源,评估仍具挑战性。我们提出ResearchRubrics,一个基于超过2800小时人工标注的标准化评估基准,包含2500+条专家撰写的细粒度评分标准,用于评估事实依据、推理严谨性和表达清晰度。我们还提出三轴复杂度框架,从概念广度、逻辑嵌套和探索深度对任务分类。同时开发了人类与模型评估协议,以衡量代理对评分标准的遵循程度。评测多个先进系统发现,即使如Gemini DR和OpenAI DR等领先模型,平均合规率也低于68%,主要因未能捕捉隐含上下文和对检索信息推理不足。结果凸显了对深度研究能力进行可靠、可扩展评估的必要性,为此我们开源ResearchRubrics(含所有提示、评分标准及评估代码),推动可信研究助手的发展。
原文摘要 · Abstract (English)
Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities, including multi-step reasoning, cross-document synthesis, and the generation of evidence-backed, long-form answers. Evaluating DR remains challenging because responses are lengthy and diverse, admit many valid solutions, and often depend on dynamic information sources. We introduce ResearchRubrics, a standardized benchmark for DR built with over 2,800+ hours of human labor that pairs realistic, domain-diverse prompts with 2,500+ expert-written, fine-grained rubrics to assess factual grounding, reasoning soundness, and clarity. We also propose a new complexity framework for categorizing DR tasks along three axes: conceptual breadth, logical nesting, and exploration. In addition, we develop human and model-based evaluation protocols that measure rubric adherence for DR agents. We evaluate several state-of-the-art DR systems and find that even leading agents like Gemini's DR and OpenAI's DR achieve under 68% average compliance with our rubrics, primarily due to missed implicit context and inadequate reasoning about retrieved information. Our results highlight the need for robust, scalable assessment of deep research capabilities, to which end we release ResearchRubrics(including all prompts, rubrics, and evaluation code) to facilitate progress toward well-justified research assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。