构建长序列桌面任务链,测试智能体持续执行能力
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks

- 用方向兼容搜索将基础任务组合成长程工作流
- 四款智能体最长任务完成率仅31%,多轮评估略有提升
- 暴露不同失败模式,适合评估智能体长期规划能力
当前计算机使用智能体的评估主要基于原子化桌面任务,但真实场景需跨目标维持状态。本文提出ChainWorld,通过方向兼容搜索将OSWorld的基础任务组合成长度为2至4的347条长程工作流,同时保留原始评估框架。评估分单轮和多轮两种:单轮中所有任务一次性给出;多轮中逐次呈现。在四款现有智能体上,最长任务链完成率仅为31%。多轮评估使三款模型表现提升,但两种协议仍具挑战性。单轮失败集中于输出精度问题,多轮失败则多表现为会话管理缺陷,如进度碎片化与后期脱离。
原文摘要 · Abstract (English)
Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap with ChainWorld, which composes atomic OSWorld tasks into long horizon desktop workloads through directional compatibility search while preserving the source evaluators. The resulting workload contains 347 chains of length two to four and compares two renderings of the same task sequence. In single turn evaluation, all tasks are presented together in one prompt. In multi turn evaluation, tasks are revealed one at a time. Across four current computer use agents, maximum chain completion is 31%. Multi turn evaluation improves completion for three models, but both protocols remain challenging. The two protocols also expose different failure profiles. Single turn failures concentrate on artifact precision, while multi turn failures more often reflect session management problems such as fragmented progress and later turn disengagement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。