arXiv:2604.08178cs.AI2026-04ACL被引 3

构建轨迹级奖励模型评估基准,测试大模型在复杂工具使用中的判断能力。

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

论文配图:Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
图 1 · 摘自论文原文
  • 设计多场景轨迹偏好数据集,模拟真实智能体决策过程。
  • 发现现有奖励模型在长序列任务中性能显著下降,尤其在错误恢复任务中。
  • 适合研究智能体对齐、奖励建模与自动化评估的学者使用。

在传统的基于人类反馈的强化学习(RLHF)中,奖励模型(RMs)是实现模型对齐的核心信号来源。随着大语言模型演化为具备自主调用工具和复杂推理能力的智能体,奖励建模范式面临前所未有的挑战,尤其是缺乏针对工具集成环境下的专门评估基准。为此,我们提出Plan-RewardBench,一个用于评估判断者在复杂工具使用场景中区分优选与干扰轨迹的轨迹级偏好基准。该基准涵盖四大代表性任务类型:(i) 安全拒绝,(ii) 工具无关性/不可用性,(iii) 复杂规划,(iv) 鲁棒错误恢复;包含经验证的正向轨迹及通过多模型自然回放、规则扰动和最小编辑LLM扰动生成的混淆难例。我们在统一的成对协议下评测代表性奖励模型(生成式、判别式及大模型作为裁判),报告不同轨迹长度和任务类别下的准确率趋势,并提供常见失败模式的诊断分析。结果表明,三类评估器均面临严峻挑战,尤其在长时序轨迹上性能急剧下降,凸显了面向智能体、轨迹级奖励建模的专用训练必要性。最终,Plan-RewardBench旨在成为可复用的评估套件与构建智能体规划偏好数据的蓝图。

原文摘要 · Abstract (English)

In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As Large Language Models evolve into agentic systems capable of autonomous tool invocation and complex reasoning, the paradigm of reward modeling faces unprecedented challenges -- most notably, the lack of benchmarks specifically designed to assess RM capabilities within tool-integrated environments. To address this gap, we present Plan-RewardBench, a trajectory-level preference benchmark designed to evaluate how well judges distinguish preferred versus distractor agent trajectories in complex tool-using scenarios. Plan-RewardBench covers four representative task families -- (i) Safety Refusal, (ii) Tool-Irrelevance / Unavailability, (iii) Complex Planning, and (iv) Robust Error Recovery -- comprising validated positive trajectories and confusable hard negatives constructed via multi-model natural rollouts, rule-based perturbations, and minimal-edit LLM perturbations. We benchmark representative RMs (generative, discriminative, and LLM-as-Judge) under a unified pairwise protocol, reporting accuracy trends across varying trajectory lengths and task categories. Furthermore, we provide diagnostic analyses of prevalent failure modes. Our results reveal that all three evaluator families face substantial challenges, with performance degrading sharply on long-horizon trajectories, underscoring the necessity for specialized training in agentic, trajectory-level reward modeling. Ultimately, Plan-RewardBench aims to serve as both a practical evaluation suite and a reusable blueprint for constructing agentic planning preference data.

奖励建模智能体对齐评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。