构建真实日常搜索任务评估基准,揭示现有搜索智能体与用户期望的差距。
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks

- 设计150个开放式日常任务,拆解为子任务并多维度评分
- 3,546条评分标准支持细粒度性能归因,结果更可解释
- 面向研究者开放数据集,推动智能体评估标准化
搜索智能体(SAs)通常利用大语言模型(LLMs)自主探索网络资源并整合信息以完成复杂的信息检索任务。现有评估基准多聚焦于特定领域任务,难以反映真实用户场景;且依赖粗粒度的任务级评分标准,限制了评估的可解释性。为此,我们提出DailyReport,一个面向日常搜索任务的开放式评估基准。该基准包含150个开放式任务和3,546条关联评分标准,涵盖真实用户广泛讨论和关注的时效性信息需求。每个任务被分解为子任务,并通过解耦维度的级联评分进行评估。结合级联性能归因与用户中心聚合,我们得出各维度的可解释得分及用户偏好分数。对17个智能体系统的评估显示,当前系统仍远未达到用户预期。为促进后续研究,数据集与代码已公开发布于https://github.com/AGI-Eval-Official/DailyReport。
原文摘要 · Abstract (English)
Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses. For SAs evaluation, prior benchmarks mainly focus on specialized tasks that are unlikely to arise in real-world user scenarios. Moreover, their reliance on coarse task-level rubrics often limits evaluation interpretability. To bridge this gap, we introduce DailyReport, an open-ended benchmark to evaluate SA capabilities on daily search tasks. It contains 150 open-ended tasks with 3,546 associated rubrics, capturing widely discussed and timely information demands of real-world users. Each task is decomposed into subtasks and evaluated with cascade rubrics across disentangled dimensions. Through cascade performance attribution and user-centric aggregation, we derive highly interpretable scores for each dimension, along with a user preference score. Our results on 17 agentic systems show that current systems still fall short of users' expectations. To facilitate future research, our dataset and code are made publicly available at https://github.com/AGI-Eval-Official/DailyReport.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。