arXiv:2512.02793cs.CV2025-12被引 4

用视频模型并行生成多视角世界视频,保持几何与运动一致性。

IC-World: In-Context Generation for Shared World Modeling

  • 激活大模型的上下文生成能力,同时处理多输入图像
  • 通过强化学习优化,生成视频在几何和运动上更一致
  • 适合需要多视角动态环境建模的研究者

基于视频的世界模型近年来受到关注,因其能合成多样且动态的视觉环境。本文聚焦共享世界建模:从一组输入图像生成多个视频,每个代表同一场景的不同相机视角。我们提出IC-World框架,利用大视频模型固有的上下文生成能力,并行生成所有输入图像对应的视频。进一步通过强化学习与新型奖励模型(组相对策略优化)微调,增强生成视频间的场景级几何一致性与物体级运动一致性。大量实验表明,IC-World在几何与运动一致性上显著优于现有方法。据我们所知,这是首个系统探索基于视频的世界模型在共享世界建模问题上的工作。

原文摘要 · Abstract (English)

Video-based world models have recently garnered increasing attention for their ability to synthesize diverse and dynamic visual environments. In this paper, we focus on shared world modeling, where a model generates multiple videos from a set of input images, each representing the same underlying world in different camera poses. We propose IC-World, a novel generation framework, enabling parallel generation for all input images via activating the inherent in-context generation capability of large video models. We further finetune IC-World via reinforcement learning, Group Relative Policy Optimization, together with two proposed novel reward models to enforce scene-level geometry consistency and object-level motion consistency among the set of generated videos. Extensive experiments demonstrate that IC-World substantially outperforms state-of-the-art methods in both geometry and motion consistency. To the best of our knowledge, this is the first work to systematically explore the shared world modeling problem with video-based world models.

世界建模视频生成一致性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。