arXiv:2509.02544cs.AIcs.CL2025-09被引 180

UI-TARS-2用多轮强化学习提升图形界面智能体的稳定性和泛化能力。

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

论文配图:UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
图 1 · 摘自论文原文
  • 构建数据飞轮与混合环境,实现可扩展的多轮强化学习训练。
  • 在多个GUI基准上性能超前代模型,最高达88.2分,游戏平均达人类60%水平。
  • 适合研究交互式智能体、自动化测试及复杂任务泛化的开发者和研究人员。

自主图形用户界面(GUI)智能体的开发面临重大挑战。尽管近期原生智能体模型通过端到端学习统一感知、推理、动作与记忆展现潜力,但在数据可扩展性、多轮强化学习、仅依赖GUI操作的局限性及环境稳定性方面仍存难题。本技术报告提出UI-TARS-2,一种以GUI为中心的原生智能体模型,通过系统性训练方法解决上述问题:数据飞轮实现可扩展数据生成,稳定化多轮强化学习框架,融合文件系统与终端的混合GUI环境,以及统一沙盒平台支持大规模部署。实证评估显示,UI-TARS-2显著优于其前身UI-TARS-1.5。在GUI基准上,分别取得Online-Mind2Web 88.2、OSWorld 47.5、WindowsAgentArena 50.6、AndroidWorld 73.3的成绩,超越Claude与OpenAI等强基线。在15款游戏组成的套件中,平均归一化得分59.8,约达人类水平的60%,在LMGame-Bench上与前沿闭源模型(如OpenAI o3)竞争力相当。该模型还能泛化至长周期信息检索与软件工程任务,展现跨任务鲁棒性。对训练动态的深入分析进一步揭示了大规模智能体强化学习中的稳定性与效率机制。结果表明,UI-TARS-2有望推动GUI智能体发展,并具备真实交互场景下的强泛化能力。

原文摘要 · Abstract (English)

The development of autonomous agents for graphical user interfaces (GUIs) presents major challenges in artificial intelligence. While recent advances in native agent models have shown promise by unifying perception, reasoning, action, and memory through end-to-end learning, open problems remain in data scalability, multi-turn reinforcement learning (RL), the limitations of GUI-only operation, and environment stability. In this technical report, we present UI-TARS-2, a native GUI-centered agent model that addresses these challenges through a systematic training methodology: a data flywheel for scalable data generation, a stabilized multi-turn RL framework, a hybrid GUI environment that integrates file systems and terminals, and a unified sandbox platform for large-scale rollouts. Empirical evaluation demonstrates that UI-TARS-2 achieves significant improvements over its predecessor UI-TARS-1.5. On GUI benchmarks, it reaches 88.2 on Online-Mind2Web, 47.5 on OSWorld, 50.6 on WindowsAgentArena, and 73.3 on AndroidWorld, outperforming strong baselines such as Claude and OpenAI agents. In game environments, it attains a mean normalized score of 59.8 across a 15-game suite-roughly 60% of human-level performance-and remains competitive with frontier proprietary models (e.g., OpenAI o3) on LMGame-Bench. Additionally, the model can generalize to long-horizon information-seeking tasks and software engineering benchmarks, highlighting its robustness across diverse agent tasks. Detailed analyses of training dynamics further provide insights into achieving stability and efficiency in large-scale agent RL. These results underscore UI-TARS-2's potential to advance the state of GUI agents and exhibit strong generalization to real-world interactive scenarios.

GUI智能体强化学习多轮交互自动化测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。