通过交互式验证环境状态,提升图形界面任务评估准确性
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

- 采用提出-验证框架,结合界面与系统状态判断任务完成
- 在321条轨迹上达86.9%准确率,显著优于现有方法
- 适用于训练GUI智能体,可生成有效奖励信号
图形用户界面任务评估旨在判断GUI代理是否成功执行用户指令。自动化评估因可作为测试时扩展和后训练的奖励信号而受到关注,但可靠评估仍具挑战性,因判断常需访问系统配置、文件数据、应用设置等环境状态,而不仅依赖执行轨迹截图。本文提出一种基于提出-验证框架的交互式奖励代理(IRA),在任务指令与执行后环境的基础上,先提出任务完成条件,再通过调用系统工具、应用工具和GUI工具进行验证。该设计将可见界面与环境状态证据在交互中融合。我们进一步构建了涵盖10类Ubuntu桌面应用的321条任务轨迹基准集GUI-RewardBench。实验表明,IRA在该基准上达到86.9%准确率,优于现有评估基线。我们将IRA应用于GUI代理的强化学习,实现34.0%的OSWorld成功率,证明其能提供有效的训练奖励信号。
原文摘要 · Abstract (English)
Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。