arXiv:2602.09022cs.CV2026-02被引 17

用强化学习让视频世界模型更准更稳地长期互动。

WorldCompass: Reinforcement Learning for Long-Horizon World Models

  • 在自回归视频生成中,按帧生成多样本并评估,提升推理效率。
  • 设计交互准确率与视觉质量双奖励函数,避免奖励欺骗。
  • 适合需要长期稳定交互的视频生成场景,如虚拟环境训练。

本文提出WorldCompass,一种针对长时序、交互式视频世界模型的强化学习后训练框架,使其能基于交互信号更准确、一致地探索世界。为有效引导模型探索,提出三项核心创新:1)片段级滚动策略:在单个目标片段生成并评估多个样本,显著提升滚动效率并提供细粒度奖励信号;2)互补奖励函数:设计交互跟随准确率与视觉质量双重奖励,提供直接监督,有效抑制奖励作弊行为;3)高效强化学习算法:采用负向感知微调策略结合多种效率优化,高效增强模型能力。在当前最先进开源世界模型WorldPlay上的实验表明,WorldCompass在多种场景下显著提升交互准确率与视觉保真度。

原文摘要 · Abstract (English)

This work presents WorldCompass, a novel Reinforcement Learning (RL) post-training framework for the long-horizon, interactive video-based world models, enabling them to explore the world more accurately and consistently based on interaction signals. To effectively "steer" the world model's exploration, we introduce three core innovations tailored to the autoregressive video generation paradigm: 1) Clip-level rollout Strategy: We generate and evaluate multiple samples at a single target clip, which significantly boosts rollout efficiency and provides fine-grained reward signals. 2) Complementary Reward Functions: We design reward functions for both interaction-following accuracy and visual quality, which provide direct supervision and effectively suppress reward-hacking behaviors. 3) Efficient RL Algorithm: We employ the negative-aware fine-tuning strategy coupled with various efficiency optimizations to efficiently and effectively enhance model capacity. Evaluations on the SoTA open-source world model, WorldPlay, demonstrate that WorldCompass significantly improves interaction accuracy and visual fidelity across various scenarios.

强化学习世界模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。