构建日常任务评测基准,检验AI能否理解自然语言并完成多样化实际任务。
AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
- 设计三类用户导向任务:流程执行、隐含指令推理、迭代优化。
- 包含104个任务、767个评分点,验证结果与人工判断一致率达80.1%。
- 适合评估通用AI Agent在真实生活场景中的实用能力,尤其对开发者有参考价值。
AI代理处理长期复杂任务的能力持续提升,在编程、深度研究和复杂问题求解中表现优异。然而,普通用户对这些高级能力的实际感知仍有限。现有评估侧重提升任务难度,却忽视了覆盖大众日常生活、工作与学习活动的多样性。为此,我们提出AgentIF-OneDay,旨在检验普通用户能否通过自然语言指令,利用AI代理完成多样化的日常任务。这些任务不仅需通过对话解决问题,还需理解多种附件类型并输出可交付的文件结果。基准涵盖三大用户中心类别:开放流程执行(评估对明确复杂流程的遵循)、隐含指令(要求从附件中推断隐性指令)、迭代优化(涉及对已有工作的修改或扩展)。采用实例级评分标准与优化评估流程,使基于LLM的验证与人工判断达到80.1%一致性,使用Gemini-3-Pro实现。该基准包含104个任务,共767个评分点。我们测试了四个主流通用AI代理,发现基于API的代理产品与基于强化学习的ChatGPT代理均处于第一梯队。领先LLM API与开源模型已内化代理能力,使团队可开发前沿代理应用。
原文摘要 · Abstract (English)
The capacity of AI agents to effectively handle tasks of increasing duration and complexity continues to grow, demonstrating exceptional performance in coding, deep research, and complex problem-solving evaluations. However, in daily scenarios, the perception of these advanced AI capabilities among general users remains limited. We argue that current evaluations prioritize increasing task difficulty without sufficiently addressing the diversity of agentic tasks necessary to cover the daily work, life, and learning activities of a broad demographic. To address this, we propose AgentIF-OneDay, aimed at determining whether general users can utilize natural language instructions and AI agents to complete a diverse array of daily tasks. These tasks require not only solving problems through dialogue but also understanding various attachment types and delivering tangible file-based results. The benchmark is structured around three user-centric categories: Open Workflow Execution, which assesses adherence to explicit and complex workflows; Latent Instruction, which requires agents to infer implicit instructions from attachments; and Iterative Refinement, which involves modifying or expanding upon ongoing work. We employ instance-level rubrics and a refined evaluation pipeline that aligns LLM-based verification with human judgment, achieving an 80.1% agreement rate using Gemini-3-Pro. AgentIF-OneDay comprises 104 tasks covering 767 scoring points. We benchmarked four leading general AI agents and found that agent products built based on APIs and ChatGPT agents based on agent RL remain in the first tier simultaneously. Leading LLM APIs and open-source models have internalized agentic capabilities, enabling AI application teams to develop cutting-edge Agent products.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。