arXiv:2608.15930cs.AIcs.CV2026-08

用演示学习提升通用图形界面智能体的可靠性和泛化能力。

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

论文配图:UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
图 1 · 摘自论文原文
  • 通过闭环环境生成数据,实现大规模任务训练与强化学习。
  • 仅需一个演示就能将长流程任务成功率从17.2%提升至35.4%。
  • 适合需要高可靠性自动化办公的开发者和研究者。

基础图形界面智能体可自动化复杂数字任务,但部署受限于稀缺且有偏的训练数据、模糊指令及不可靠执行。日常流程依赖用户特定工具与隐含规范,未明示指令易导致运行结果差异。我们提出UI-Mate,一种融合环境感知训练体系与上下文示范学习的通用GUI智能体。其三大贡献包括:可扩展的环境感知训练栈——通过统一的任务验证包,在大规模并行环境中自动完成任务生成、环境构建、执行回放、过滤、能力平衡、监督微调与在线强化学习;上下文示范学习机制——将多模态示范转化为灵活的子任务级工作流,能跟随已演示步骤并基于实时界面重规划;以及OSWorkerBench基准与洞察——包含41个应用中的100个长时程办公任务,支持仅指令与示范引导两种评估方式。该基准将33个自示范任务(由强智能体成功回放构建)与45个变示范任务(由人类记录的非完全相同任务构建)分离。实验表明,UI-Mate-27B在通用计算机使用基准上达到新开放权重最佳性能:在OSWorld-Verified上得分为77.0%,在WindowsAgentArena上为66.2%。在OSWorkerBench上,严格成功率达41.0%,进度达76.9%,较Qwen3.6-27B基线分别提升17.7与24.5点。在33个自示范子集上,仅一个示范即可使严格成功率从17.2%升至35.4%,进度从67.9%升至81.1%,显著提升长周期任务可靠性。

原文摘要 · Abstract (English)

Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce arbitrary variations across runs. We present UI-Mate, a foundation GUI agent that integrates an environment-grounded training stack with in-context demonstration learning. UI-Mate makes three contributions: A Scalable Environment-Grounded Training Stack: A closed-loop data engine automates task generation, environment construction, rollout, filtering, capability balancing, SFT, and online RL across massively parallel environments via unified task-verifier bundles. In-Context Demonstration Learning: A mechanism that transforms multimodal demonstrations into flexible subtask-level workflows, follows relevant demonstrated steps, and re-plans from the live interface. OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation. Its demonstration resources separate a 33-task self-demo setting, built from successful strong-agent rollouts of the same targets, from a 45-task variant-demo setting, built from human recordings of related but non-identical tasks. Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, scoring 77.0% on OSWorld-Verified and 66.2% on WindowsAgentArena. On OSWorkerBench, it reaches 41.0% strict success and 76.9% progress, outperforming its Qwen3.6-27B base by 17.7 and 24.5 points. On the 33-task self-demo subset, one demonstration raises strict success from 17.2% to 35.4% and progress from 67.9% to 81.1%, substantially improving long-horizon reliability. Project page: https://ui-mate.github.io.

GUI智能体示范学习自动化办公强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。