用自演化数据与锚点奖励提升移动界面智能体训练效率
GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

- 自演化生成多环境,自动标注成功状态的界面元素作锚点
- 在安卓世界和新构建基准上均达90%以上验证准确率
- 适合需要高效稳定训练的GUI智能体研究者使用
基于视觉语言模型的图形用户界面(GUI)智能体可显著受益于在线强化学习(RL)。然而,其训练受限于两大核心问题:现有数据合成方法依赖特定环境,难以生成多样化数据;现有评估器或存在扩展性差,或提供不准确、不可靠的奖励信号。为此,我们提出GSAR(目标状态锚定奖励)框架,支持可扩展的任务生成并提供可靠奖励信号,实现稳定高效的策略优化。该方法采用自演化数据合成,通过任务执行生成多个环境,产生多样化任务与目标状态;同时引入状态锚定机制,自动标注成功目标状态中的相关界面元素作为参考锚点。在强化学习训练中,这些参考锚点提供精确且可扩展的奖励信号,显著提升训练效率。大量实验表明,本框架在离线轨迹验证中准确率超过90%,表现最接近规则基方法。基于该奖励框架训练的智能体在AndroidWorld及自建基准上均表现优异,确立了一种可扩展的GUI智能体训练范式。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific environments and struggle to generate diverse data, while existing evaluators either suffer from limited scalability or provide inaccurate and unreliable reward signals. To overcome these challenges, we introduce GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization. Our approach features self-evolving data synthesis, which produces multiple environments through task execution and generates diverse tasks and goal states. Complementing this, a state-anchor mechanism automatically annotates task-relevant UI elements in successful goal states as reference anchors. During RL training, these reference anchors provide accurate, scalable reward signals that substantially enhance efficiency. Extensive evaluations demonstrate that our framework achieves over 90% accuracy on offline trajectory verification and performs closest to rule-based methods. Furthermore, agents trained using our reward framework exhibit strong performance on both AndroidWorld and our constructed benchmark, establishing a scalable approach for GUI agent training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。