提出可提供逐步反馈的GUI任务奖励模型,提升长序列操作成功率。
GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- 用人类标注与GPT-4o生成推理构建5.2万条交互数据,训练过程奖励模型。
- 在线任务中提升成功率7.7点,作为验证器时提升5.1点,效果显著。
- 适用于在线强化学习与离线推理,通用性强,适合复杂GUI自动化场景。
面向长序列图形用户界面(GUI)任务的自主智能体受限于稀疏奖励与难以解决的信用分配问题。为此,我们提出GUI-Shepherd,一种过程奖励模型,可提供密集的、逐步反馈以引导智能体。该模型基于包含52,000次交互的多样化大规模数据集进行训练,数据包含人工标注评分和GPT-4o生成的推理过程,使其既能作为强化学习训练中的奖励提供者,也能在推理阶段充当验证器。据我们所知,这是首个在从在线长时程任务到离线单步预测等多种场景下系统研究过程监督的尝试。在在线AndroidWorld基准测试中,通过多轮在线PPO优化,成功率达7.7个百分点的提升,显著优于基于结果奖励模型的竞品。作为推理验证器使用时,性能提升达5.1个百分点。该优势在离线AndroidControl基准上也得到验证,分别带来2.2和4.3个百分点的提升。综合结果表明,高保真过程监督对构建更强大的GUI智能体至关重要,并提供了一种可泛化的解决方案。
原文摘要 · Abstract (English)
Autonomous agents for long-sequence Graphical User Interface tasks are hindered by sparse rewards and the intractable credit assignment problem. To address these challenges, we introduce GUI-Shepherd, a Process Reward Model that provides dense, step-by-step feedback to guide agents. GUI-Shepherd is trained on a diverse large-scale data set of $52$k interactions that features human-annotated scores and GPT-4o generated rationales, enabling it to serve both as a reward provider for RL training and as a verifier for inference. As far as we know, we are the first to conduct a systematic study of process supervision in GUI agents, across diverse settings from online long-horizon tasks to offline single-step prediction. On the online AndroidWorld benchmark, GUI-Shepherd improves success rate by $7.7$ points via multi-turn online PPO, significantly outperforming Outcome Reward Model based competitors. When used as an inference verifier, it brings $5.1$ points improvements. The benefits generalize to the offline AndroidControl benchmark, with gains of $2.2$ points as a reward provider and $4.3$ points as a verifier. Collectively, our results establish that high-fidelity process supervision is critical for building more capable GUI agents and present a generalizable solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。