arXiv:2509.23958cs.CV2025-09被引 11

用逆动力学模型反推动作,让世界模型更准确跟从人类指令。

Reinforcement Learning with Inverse Rewards for World Model Post-training

  • 通过逆动力学模型从生成视频中恢复输入动作,构建可验证奖励信号。
  • 在自回归与扩散模型上实现5-10%动作跟随提升,视觉质量最高提升10%。
  • 适合希望提升视频世界模型动作一致性与可控性的研究人员。

世界模型能模拟动态环境,使智能体与多模态输入交互。尽管近期进展提升了视频世界模型的视觉质量和时间一致性,其对人类指定动作的准确建模能力仍待深入探索。强化学习为直接改进预训练模型的动作跟随性能提供了可能,前提是可定义合适的奖励函数。然而,将强化学习后训练方法迁移至世界模型不切实际,因大规模偏好标注成本过高,且基于规则的视频验证器难以构建。为此,我们提出逆奖励强化学习(RLIR),一种后训练框架,通过逆动力学模型从生成视频中恢复输入动作,从而获得可验证的奖励信号。通过将高维视频模态映射到低维动作空间,RLIR利用组相对策略优化提供客观奖励进行优化。在自回归与扩散范式上的实验表明,动作跟随性能提升5-10%,视觉质量最高提升10%,人类偏好评分更高,确立了RLIR作为首个专门用于增强视频世界模型动作跟随能力的后训练方法。

原文摘要 · Abstract (English)

World models simulate dynamic environments, enabling agents to interact with diverse input modalities. Although recent advances have improved the visual quality and temporal consistency of video world models, their ability of accurately modeling human-specified actions remains under-explored. Reinforcement learning presents a promising approach for directly improving the suboptimal action-following capability of pre-trained models, assuming that an appropriate reward function can be defined. However, transferring reinforcement learning post-training methods to world model is impractical due to the prohibitive cost of large-scale preference annotations and the infeasibility of constructing rule-based video verifiers. To address this gap, we propose Reinforcement Learning with Inverse Rewards (RLIR), a post-training framework that derives verifiable reward signals by recovering input actions from generated videos using an Inverse Dynamics Model. By mapping high-dimensional video modality to a low-dimensional action space, RLIR provides an objective and verifiable reward for optimization via Group Relative Policy Optimization. Experiments across autoregressive and diffusion paradigms demonstrate 5-10% gains in action-following, up to 10% improvements in visual quality, and higher human preference scores, establishing RLIR as the first post-training method specifically designed to enhance action-following in video world models.

世界模型强化学习动作跟随视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。