arXiv:2506.21252cs.CLcs.AI2025-06ACL被引 19

为多模态智能体设计统一奖励评估基准,解决反馈难题。

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents

  • 构建覆盖感知、规划、安全的7个真实场景评估体系
  • 支持任务每一步的奖励打分,实现精细化性能分析
  • 适合研究智能体训练与奖励建模的学者使用

随着多模态大语言模型(MLLMs)的发展,多模态智能体在网页导航和具身智能等真实任务中展现出潜力。然而,由于缺乏外部反馈,这些智能体在自我修正和泛化能力上存在局限。利用奖励模型作为外部反馈是一种有前景的解决方案,但目前尚无明确标准来选择适合的奖励模型。因此,亟需建立面向智能体的奖励评估基准。为此,我们提出 Agent-RewardBench,一个用于评估 MLLMs 奖励建模能力的基准。该基准具有三大特点:(1)涵盖感知、规划、安全三个维度,包含7个真实世界智能体场景;(2)支持步骤级奖励评估,可对任务执行过程中的每一步进行性能分析,提供更细粒度的评估视角;(3)难度适中且数据质量高,从10个多样化模型中精心采样,通过难度控制保持任务挑战性,并经人工验证确保数据完整性。实验表明,即使是当前最先进的多模态模型,在奖励建模方面表现仍有限,凸显了专门训练的必要性。代码已开源。

原文摘要 · Abstract (English)

As Multimodal Large Language Models (MLLMs) advance, multimodal agents show promise in real-world tasks like web navigation and embodied intelligence. However, due to limitations in a lack of external feedback, these agents struggle with self-correction and generalization. A promising approach is to use reward models as external feedback, but there is no clear on how to select reward models for agents. Thus, there is an urgent need to build a reward bench targeted at agents. To address these challenges, we propose Agent-RewardBench, a benchmark designed to evaluate reward modeling ability in MLLMs. The benchmark is characterized by three key features: (1) Multiple dimensions and real-world agent scenarios evaluation. It covers perception, planning, and safety with 7 scenarios; (2) Step-level reward evaluation. It allows for the assessment of agent capabilities at the individual steps of a task, providing a more granular view of performance during the planning process; and (3) Appropriately difficulty and high-quality. We carefully sample from 10 diverse models, difficulty control to maintain task challenges, and manual verification to ensure the integrity of the data. Experiments demonstrate that even state-of-the-art multimodal models show limited performance, highlighting the need for specialized training in agent reward modeling. Code is available at github.

智能体奖励建模多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。