arXiv:2606.03103cs.AI2026-06被引 2

评测桌面智能体在复杂专业工作流中的协作能力,强调长期任务与实时互动。

DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration

论文配图:DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
图 1 · 摘自论文原文
  • 构建多层级难度的长时序桌面任务基准,覆盖设计、视频、音频等专业软件
  • 18个智能体在538项任务中表现有限,交互任务准确率仅27.6%
  • 首次系统建模人机协同全流程,支持主动追问与中途干预

真实世界的专业桌面工作流在创意与工程软件中持续时间长,常需人机协同:智能体主动获取信息,用户则在过程中提供指令、澄清或修正。然而现有桌面GUI基准大多简化为短时任务,且所有指令提前给出。为此,我们提出DeskCraft,一个面向长时序创意与工程工作流的桌面GUI基准,支持主动的人机协作。DeskCraft采用多层级难度分类,任务超过50步执行,涵盖设计、视频、音频、3D创作等专业软件。同时,它形式化了人机协作流程,包含中段交互(智能体不确定时主动询问,用户执行中中断)与后段交互(智能体完成时用户反馈),完整覆盖现实协作模式。我们在538个任务上评估18个专有及开源智能体,发现GPT-5.4在标准任务上达31.6%,交互任务仅27.6%。深入分析显示,长期任务交付与主动澄清仍存在明显缺陷。所有评估代码、任务和数据将开源至https://github.com/mrwwk/DeskCraft。

原文摘要 · Abstract (English)

Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents proactively seek necessary information and users provide additional instructions, clarifications, feedback, or corrections as the task progresses. Yet existing desktop GUI benchmarks mostly reduce this setting to short, simplified tasks with all user instructions provided upfront. To address this issue, we introduce DeskCraft, a desktop GUI benchmark targeting long horizon creative and engineering workflows and proactive human-agent collaboration. DeskCraft organizes tasks into a multilevel difficulty taxonomy, with long horizon tasks requiring over 50 execution steps, and covers professional creative software across design, video, audio, and 3D creation. Furthermore, DeskCraft formalizes human-agent collaboration into an interaction protocol covering mid-turn and post-turn exchanges. Mid-turn interaction captures both agent-initiated clarification under uncertainty and user-initiated interruption during execution, while post-turn interaction accommodates user-driven feedback after the agent signals completion, together spanning the full space of realistic collaboration patterns. We evaluate 18 proprietary and open source agents on 538 tasks and find that GPT-5.4 reaches 31.6% on standard tasks and 27.6% on interactive tasks. Further analyses reveal persistent failures in long horizon workflow delivery and proactive clarification. We will open-source all evaluation codes, tasks, and data at https://github.com/mrwwk/DeskCraft.

人机协作桌面智能体长时序任务基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。