arXiv:2505.15277cs.CL2025-05NeurIPS被引 29

首个面向网页导航的步骤级奖励模型,提升智能体决策效率与成本效益。

Web-Shepherd: Advancing PRMs for Reinforcing Web Agents

  • 构建步骤级奖励模型Web-Shepherd,支持训练与推理阶段评估
  • 在基准测试中比GPT-4o高30分准确率,性能更优且成本更低
  • 适合需要高效、低成本网页自动化部署的研究与开发者

网页导航是需长序列决策的独特领域,可自动化重复性任务,但传统多模态大模型难以胜任。现有工作多依赖大模型作奖励模型,限制实际部署。本文提出首个过程奖励模型(PRM)Web-Shepherd,实现对网页导航轨迹的步骤级评估。首先构建包含4万条步骤级偏好对的WebPRM Collection数据集,并附带跨领域、多难度的检查清单。其次引入首个用于评估PRMs的元评测基准WebRewardBench。实验显示,Web-Shepherd在WebRewardBench上较GPT-4o高出约30分准确率;在WebArena-lite测试中,使用GPT-4o-mini为策略、Web-Shepherd为验证器时,性能优于用GPT-4o-mini作为验证器,提升10.9分,同时节省10倍成本。模型、数据集与代码已公开。

原文摘要 · Abstract (English)

Web navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimodal large language model (MLLM) tasks. Yet, specialized reward models for web navigation that can be utilized during both training and test-time have been absent until now. Despite the importance of speed and cost-effectiveness, prior works have utilized MLLMs as reward models, which poses significant constraints for real-world deployment. To address this, in this work, we propose the first process reward model (PRM) called Web-Shepherd which could assess web navigation trajectories in a step-level. To achieve this, we first construct the WebPRM Collection, a large-scale dataset with 40K step-level preference pairs and annotated checklists spanning diverse domains and difficulty levels. Next, we also introduce the WebRewardBench, the first meta-evaluation benchmark for evaluating PRMs. In our experiments, we observe that our Web-Shepherd achieves about 30 points better accuracy compared to using GPT-4o on WebRewardBench. Furthermore, when testing on WebArena-lite by using GPT-4o-mini as the policy and Web-Shepherd as the verifier, we achieve 10.9 points better performance, in 10 less cost compared to using GPT-4o-mini as the verifier. Our model, dataset, and code are publicly available at LINK.

网页导航奖励模型智能体高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。