用智能体式奖励机制,让机器人世界模型更敢试、更准验。
Reward as An Agent for Embodied World Models

- 设计可主动评估行为的智能体奖励,防止奖励作弊。
- 通过动态感知的多样化采样,拓展动作空间探索范围。
- 在复杂动态中实现更可靠、更丰富的行为学习,适合具身智能研究者。
尽管强化学习已成为优化世界模型的有力工具,但现有方法大多依赖靠近训练分布的保守采样,限制了探索能力、行为多样性及动态发现。本文挑战这一保守范式,认为核心问题不在于探索本身,而在于缺乏可靠的验证策略以支撑更大范围探索。缺乏验证时,扩展探索极易导致奖励欺骗——策略利用不完善的奖励信号却未实现真实改进。为此,我们在具身世界模型中构建验证框架:引入「奖励作为智能体」(Reward as an Agent),由一个能主动评估生成行为的智能体提供稳健奖励信号,有效抑制分布偏移下的奖励欺骗。同时提出基于动态感知的滚动多样化方法 DynDiff-GRPO,显式扩展动作空间探索,丰富轨迹多样性,扩大状态-动作覆盖范围,推动超越传统保守采样范式的具身行为演化。二者结合后,在多个开源世界模型上显著提升准确率,证明只要建立可靠验证基础,更广泛的探索可实现高效且稳定的学习。
原文摘要 · Abstract (English)
While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distribution, limiting exploration, behavioral diversity, and richer dynamic discovery. In this work, we challenge this conservative paradigm. We argue that the core limitation is not exploration itself, but the lack of reliable verification strategies to support broader exploration. Without reliable verification, expanded exploration becomes highly susceptible to reward hacking, where policies exploit imperfect rewards without achieving genuine improvement. To evaluate this motivation, we instantiate our method in embodied world models, where physical plausibility, and task completion provide a rigorous testbed for scalable RL under complex dynamics. On the verification side, we introduce Reward as an Agent, an agentic reward framework that actively evaluates generated behaviors to provide robust reward signals and mitigate reward hacking under distribution shifts. On the exploration side, we introduce Dynamic-Aware Rollout Diversification through DynDiff-GRPO, which explicitly expands action-space exploration to diversify trajectories, broaden state-action coverage, and encourage richer embodied behaviors beyond conservative rollout regimes. By unifying Reward as an Agent with DynDiff-GRPO, we enable RL on a more reliable reward foundation with substantially diversified sampling, effectively mitigating reward hacking while yielding significant accuracy gains across multiple open-source world models, thereby demonstrating that broader exploration can scale successfully when grounded in robust verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。