搭建平台实现深度研究智能体的细粒度人工评估与对比
Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
- 提供双智能体并列展示与中间步骤追踪的评估框架
- 17名标注者对3个智能体进行偏好打分,支持逐阶段反馈
- 内置基础架构可快速集成大模型,适合研究者测试新方法
评估能够自主上网搜索、分析信息并生成报告的深度研究智能体仍面临重大挑战,尤其是在评估长篇报告和提供中间步骤详细反馈方面。为解决这一问题,我们推出了Deep Research Comparator平台,该平台提供深度研究智能体托管、并排比较、细粒度人工反馈收集及排名计算的完整框架。针对用户查询,平台同时展示两个不同智能体生成的最终报告及其生成过程中的中间步骤。标注者可基于并列比较评估报告整体质量,并分别对中间步骤或最终报告中的特定文本段落提供详细反馈。此外,我们开发了Simple Deepresearch——一个端到端的智能体基线框架,可简化多种大语言模型的集成,使其快速转化为深度研究智能体以供评估。为验证平台在智能体开发中的实用性,我们已收集来自17名标注者的三组真实用户偏好数据。平台演示视频见https://www.youtube.com/watch?v=g4d2dnbdseg。
原文摘要 · Abstract (English)
Effectively evaluating deep research agents that autonomously search the web, analyze information, and generate reports remains a major challenge, particularly when it comes to assessing long reports and giving detailed feedback on their intermediate steps. To address these gaps, we introduce Deep Research Comparator, a platform that offers a holistic framework for deep research agent hosting, side-by-side comparison, fine-grained human feedback collection, and ranking calculation. Given a user query, our platform displays the final reports from two different agents along with their intermediate steps during generation. Annotators can evaluate the overall quality of final reports based on side-by-side comparison, and also provide detailed feedback separately by assessing intermediate steps or specific text spans within the final report. Furthermore, we develop Simple Deepresearch, an end-to-end agent scaffold. This scaffold serves as a baseline that facilitates the easy integration of various large language models to transform them into deep research agents for evaluation. To demonstrate the platform's utility for deep research agent development, we have collected real user preference data from 17 annotators on three deep research agents. A demo video of our platform can be found at https://www.youtube.com/watch?v=g4d2dnbdseg.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。