arXiv:2601.12428cs.ROcs.CV2026-01被引 7

让机器人世界模型更真实,兼顾物理、逻辑和视觉效果。

ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models

  • 构建23.5万条视频偏好数据,训练多维度奖励模型。
  • 生成视频在物理真实性、任务逻辑上显著优于现有方法。
  • 适合需要真实交互的机器人模拟与决策研究者。

近期基于视频的世界模型在机器人学习中受到关注,但现有方法多聚焦视觉生成质量,忽视物理真实性、动态一致性及任务逻辑,尤其在高接触操作任务中表现受限。为此,我们提出ReWorld框架,通过强化学习对齐视频驱动的具身世界模型与物理现实性、任务完成能力、具身合理性及视觉质量。首先构建大规模(约23.5万条)视频偏好数据集,并训练分层奖励模型以捕捉符合人类偏好的多维度奖励。进一步提出高效对齐算法,使用该奖励通过类PPO方法微调流模型世界模型。大量实验与理论分析表明,ReWorld显著提升生成轨迹的物理真实性、逻辑连贯性、具身合理性和视觉质量,优于先前方法。

原文摘要 · Abstract (English)

Recently, video-based world models that learn to simulate the dynamics have gained increasing attention in robot learning. However, current approaches primarily emphasize visual generative quality while overlooking physical fidelity, dynamic consistency, and task logic, especially for contact-rich manipulation tasks, which limits their applicability to downstream tasks. To this end, we introduce ReWorld, a framework aimed to employ reinforcement learning to align the video-based embodied world models with physical realism, task completion capability, embodiment plausibility and visual quality. Specifically, we first construct a large-scale (~235K) video preference dataset and employ it to train a hierarchical reward model designed to capture multi-dimensional reward consistent with human preferences. We further propose a practical alignment algorithm that post-trains flow-based world models using this reward through a computationally efficient PPO-style algorithm. Comprehensive experiments and theoretical analysis demonstrate that ReWorld significantly improves the physical fidelity, logical coherence, embodiment and visual quality of generated rollouts, outperforming previous methods.

世界模型机器人学习多维奖励具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。