用主动交互提升GUI智能体的验证准确率
Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
- 设计逐层探测机制,结合轨迹证据与环境探查
- 在两个基准上验证准确率显著提升
- 适合需要高可靠性的自动化测试场景
基于可验证奖励的强化学习为持续提升GUI智能体提供了新路径,但现有奖励建模方法存在互补性局限。规则方法扩展性差,难以应对开放任务;大模型作为裁判的方法虽可扩展,却被动且受限于状态可观测性,关键证据常存在于轨迹之外的隐含环境状态中。近期主动交互方法缓解了可观测性问题,但过度依赖探查而忽视直接轨迹证据,导致验证效率低下。为此,我们提出一种基于轨迹的交互验证范式,引入VAGEN框架:一个由渐进式验证机制驱动的工具增强型验证智能体,遵循由表及里、由低成本到高成本的设计理念,主动地端到端提取轨迹证据并探查环境状态。在OSWorld-Verified和AndroidWorld基准上的实验表明,VAGEN显著提升了评估准确率,同时具备良好的性能-效率权衡。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) provides a promising pathway for continuously advancing GUI agents, yet existing reward modeling paradigms face complementary limitations. Rule-based methods suffer from poor scalability and cannot handle open-ended tasks. LLM-as-a-Judge methods enable scalable trajectory verification but remain passive and are constrained by partial state observability, since key evidence often resides in latent environment states beyond the trajectory. Recent active environment interaction methods mitigate observability issues but tend to over-rely on probing while under-utilizing direct trajectory evidence, leading to verification inefficiency. To address these challenges, we advocate a trajectory-grounded interactive verification paradigm. We introduce VAGEN, a framework that employs a tool-augmented verifier agent governed by a Progressive Verification Mechanism, which follows a surface-to-latent and cheap-to-expensive design philosophy to extract trajectory evidence and probe environment states in a proactive end-to-end manner. Experiments on OSWorld-Verified and AndroidWorld benchmarks demonstrate that VAGEN significantly improves evaluation accuracy with a favorable performance-efficiency trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。