让AI通过失败经验自我进化,高效完成安卓界面自动化操作。
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
- 用拒绝式微调实现数据与模型的自主协同演化。
- 在安卓世界数据集上达81.0%成功率,超越人类水平。
- 无需人工标注,适合长期复杂任务的自动化研究者。
随着多模态大语言模型的发展,自主移动GUI代理受到越来越多关注。然而,现有方法在长周期任务中仍存在学习失败轨迹效率低、稀疏奖励下信用分配模糊等问题。为此,我们提出UI-Voyager,一种两阶段自演化移动GUI代理。第一阶段采用拒绝式微调(RFT),实现数据与模型的完全自主协同演化;第二阶段引入组相对自蒸馏(GRSD),识别群体回放中的关键分支点,从成功轨迹构建密集步骤级监督以修正失败轨迹。在AndroidWorld上的实验表明,4B模型在Pass@1指标上达到81.0%的成功率,显著优于多个近期基线,并超过人类表现。消融实验与案例分析进一步验证了GRSD的有效性。该方法实现了无需昂贵人工标注的高效、自演化、高性能移动GUI自动化。
原文摘要 · Abstract (English)
Autonomous mobile GUI agents have attracted increasing attention along with the advancement of Multimodal Large Language Models (MLLMs). However, existing methods still suffer from inefficient learning from failed trajectories and ambiguous credit assignment under sparse rewards for long-horizon GUI tasks. To that end, we propose UI-Voyager, a novel two-stage self-evolving mobile GUI agent. In the first stage, we employ Rejection Fine-Tuning (RFT), which enables the continuous co-evolution of data and models in a fully autonomous loop. The second stage introduces Group Relative Self-Distillation (GRSD), which identifies critical fork points in group rollouts and constructs dense step-level supervision from successful trajectories to correct failed ones. Extensive experiments on AndroidWorld show that our 4B model achieves an 81.0% Pass@1 success rate, outperforming numerous recent baselines and exceeding human-level performance. Ablation and case studies further verify the effectiveness of GRSD. Our method represents a significant leap toward efficient, self-evolving, and high-performance mobile GUI automation without expensive manual data annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。