arXiv:2604.11201cs.CLcs.AI2026-04被引 6

评测能融合视觉、搜索与编码的统一数字代理,发现当前系统成功率仅45.1%。

CocoaBench: Evaluating Unified Digital Agents in the Wild

  • 构建人类设计的长时序任务,要求代理灵活组合多能力。
  • 现有最佳代理在新基准上成功率仅45.1%,远未可靠。
  • 适合研究通用智能体、推理规划与工具使用的研究者。

大型语言模型代理在软件工程、深度研究、GUI自动化等任务中表现优异,近期架构正将这些能力整合为统一系统。然而,多数评估仍孤立测试各项能力,难以覆盖需多技能协同的真实场景。本文提出CocoaBench,一个基于人类设计的长时序任务基准,要求代理灵活组合视觉、搜索与编码能力。任务仅通过指令和最终输出的自动评估函数定义,支持跨多种代理架构的可靠、可扩展评估。同时推出CocoaAgent,一种轻量级共享框架,用于模型主干的可控对比。实验显示,当前代理在CocoaBench上表现仍不理想,最优系统成功率为45.1%。分析表明,推理规划、工具调用与执行、视觉定位等方面仍有巨大提升空间。

原文摘要 · Abstract (English)

LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly integrating these capabilities into unified systems. Yet, most evaluations still test these capabilities in isolation, which leaves a gap for more diverse use cases that require agents to combine different capabilities. We introduce CocoaBench, a benchmark for unified digital agents built from human-designed, long-horizon tasks that require flexible composition of vision, search, and coding. Tasks are specified only by an instruction and an automatic evaluation function over the final output, enabling reliable and scalable evaluation across diverse agent infrastructures. We also present CocoaAgent, a lightweight shared scaffold for controlled comparison across model backbones. Experiments show that current agents remain far from reliable on CocoaBench, with the best evaluated system achieving only 45.1% success rate. Our analysis further points to substantial room for improvement in reasoning and planning, tool use and execution, and visual grounding.

数字代理多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。