arXiv:2505.21964cs.HCcs.CL2025-05中稿 · ICML被引 1

让电脑操作智能体自动进化知识,提升执行成功率。

UI-Evol: Automatic Knowledge Evolving for Computer Use Agents

  • 通过真实交互回溯动作序列,自动提取有效操作路径。
  • 90%正确知识仅达41%执行成功率,引入新模块后显著提升。
  • 适合研究智能体可靠性与自动化知识演化的开发者。

外部知识在计算机使用智能体的发展中起关键作用。我们识别出一个关键的知识-执行差距:检索到的知识难以转化为有效的现实任务执行。分析显示,即使知识正确率高达90%,执行成功率也仅达41%。为弥合这一差距,我们提出UI-Evol,一个可即插即用的自主GUI知识演化模块。UI-Evol包含两个阶段:重溯阶段从智能体-环境的真实交互中提取忠实的动作序列;批判阶段则通过对比这些序列与外部参考,优化现有知识。我们在OSWorld基准上对最先进的Agent S2进行了全面实验。结果表明,UI-Evol不仅显著提升任务性能,还解决了此前被忽视的高行为方差问题,使智能体在计算机使用任务中表现更优,可靠性大幅提升。

原文摘要 · Abstract (English)

External knowledge has played a crucial role in the recent development of computer use agents. We identify a critical knowledge-execution gap: retrieved knowledge often fails to translate into effective real-world task execution. Our analysis shows even 90% correct knowledge yields only 41% execution success rate. To bridge this gap, we propose UI-Evol, a plug-and-play module for autonomous GUI knowledge evolution. UI-Evol consists of two stages: a Retrace Stage that extracts faithful objective action sequences from actual agent-environment interactions, and a Critique Stage that refines existing knowledge by comparing these sequences against external references. We conduct comprehensive experiments on the OSWorld benchmark with the state-of-the-art Agent S2. Our results demonstrate that UI-Evol not only significantly boosts task performance but also addresses a previously overlooked issue of high behavioral standard deviation in computer use agents, leading to superior performance on computer use tasks and substantially improved agent reliability.

智能体知识演化可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。