arXiv:2512.03556cs.ROcs.CV2025-12被引 2

用世界模型生成通用奖励,提升机器人在陌生环境中的泛化能力。

RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL

  • 基于世界模型构建内在奖励机制,无需人工设计奖励函数。
  • 在域外场景下平均性能比基线提升37.5%。
  • 适合需要跨场景泛化的机器人强化学习研究者。

实现具身智能体的泛化能力仍是关键挑战。传统策略学习范式,包括模仿学习(IL)和强化学习(RL),难以在多样场景中保持泛化性。IL策略常过拟合特定专家轨迹,而RL则缺乏统一、通用的奖励信号以支持多场景泛化。我们认为世界模型可作为通用环境代理来解决此问题。然而,现有世界模型主要聚焦于观测预测,仍依赖任务特定的手工奖励函数,无法提供真正通用的训练环境。为此,我们提出RoboScape-R框架,利用世界模型作为强化学习中具身环境的多功能代理。引入一种基于世界模型的通用奖励机制,从模型对真实世界状态转移动态的内在理解中生成“内生”奖励。大量实验表明,RoboScape-R有效克服了传统RL方法的局限,提供了高效且通用的训练环境,显著提升具身策略的泛化能力。该方法为将世界模型用于在线训练提供了关键洞见,在域外场景下平均性能比基线提升37.5%。

原文摘要 · Abstract (English)

Achieving generalizable embodied policies remains a key challenge. Traditional policy learning paradigms, including both Imitation Learning (IL) and Reinforcement Learning (RL), struggle to cultivate generalizability across diverse scenarios. While IL policies often overfit to specific expert trajectories, RL suffers from the inherent lack of a unified and general reward signal necessary for effective multi-scene generalization. We posit that the world model is uniquely capable of serving as a universal environment proxy to address this limitation. However, current world models primarily focus on their ability to predict observations and still rely on task-specific, handcrafted reward functions, thereby failing to provide a truly general training environment. Toward this problem, we propose RoboScape-R, a framework leveraging the world model to serve as a versatile, general-purpose proxy for the embodied environment within the RL paradigm. We introduce a novel world model-based general reward mechanism that generates ''endogenous'' rewards derived from the model's intrinsic understanding of real-world state transition dynamics. Extensive experiments demonstrate that RoboScape-R effectively addresses the limitations of traditional RL methods by providing an efficient and general training environment that substantially enhances the generalization capability of embodied policies. Our approach offers critical insights into utilizing the world model as an online training strategy and achieves an average 37.5% performance improvement over baselines under out-of-domain scenarios.

机器人强化学习世界模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。