arXiv:2607.28227cs.AIcs.CV2026-07被引 3

Qwen-UI-Agent让电脑手机网页都能自动操作,还能自己改进。

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

论文配图:Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
图 1 · 摘自论文原文
  • 统一界面与命令行操作,一次模型输出多步动作。
  • 移动端准确率达97.5%,跨平台任务完成率超80%。
  • 可主动发起服务并持续自我优化,适合真实场景使用。

GUI代理有望成为现有数字设备的通用执行者。为推动其实现真实世界应用,我们提出面向真实设备、跨平台协同、结合图形界面与命令行操作、完成长程任务、主动启动服务并自主提升能力的愿景。基于此,我们推出Qwen-UI-Agent——一个覆盖移动、计算机、网页及DeepSearch环境的实时世界中心型基础GUI代理。其融合多样沙盒环境与大规模真实移动端运行时,统一动作空间在单次模型调用中交织执行GUI操作与CLI指令,并生成批处理动作。采用类似AutoResearch的数据飞轮机制,通过代理自构建任务与环境、诊断失败并规划迭代。在线强化学习支持训练超过100步的轨迹,超过10,000个并发环境加速数据采集。轻量级封装层支持主动服务触发与跨移动/计算机的状态化工作流。在广泛评估中,Qwen-UI-Agent在移动端基准上达到领先水平,在计算机与浏览器任务上相较前沿模型(如Opus 4.8、Gemini 3.1 Pro、GPT-5.6 Sol)表现优异:移动端在MobileWorld达82.1%,MobileWorld-Real为92.2%,AndroidDaily为97.5%;计算机端在OSWorld-Verified达79.5%,OSWorld-v2部分进度得分40.0%;浏览器与GUI定位任务中,WebArena为73.6%,ScreenSpot-Pro为81.5%。

原文摘要 · Abstract (English)

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.

GUI代理自动化多模态自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。