arXiv:2603.25420cs.CV2026-03

让多视角视频生成保持外观一致,支持机器人迁移学习。

VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents

  • 用共享4D隐空间统一多视角,解决跨视角不一致问题。
  • 在单视角上达到顶尖性能,首次实现多视角物理与风格一致生成。
  • 适合做机器人世界随机化、多相机仿真训练的研究者使用。

视频到视频(V2V)翻译的进展使具身智能演示的逼真重演成为可能,使预训练机器人策略可在无需额外数据采集的情况下迁移到新环境。然而,以往方法仅支持单视角,而具身智能任务通常由多个同步摄像头捕获以支持策略学习。若对各视角独立应用单视角模型,会导致跨视角外观不一致;标准Transformer架构因跨视角注意力呈二次增长,难以扩展至多视角场景。我们提出VideoWeaver,首个多模态多视角V2V翻译框架。该模型初始为单视角流式V2V模型,通过将所有视角锚定于基于前馈空间基础模型Pi3生成的共享4D隐空间,实现视图一致性,即使在大基线和动态相机运动下也有效。为突破固定摄像头数量限制,我们在不同扩散时间步训练各视角,使模型学习联合与条件视图分布,从而实现基于已有视角的自回归新视角合成。实验表明,其在单视角基准上表现优于或相当主流方法,并首次实现物理与风格一致的多视角生成,涵盖机器人学习中关键的自拍视角与异构相机设置。

原文摘要 · Abstract (English)

Recent progress in video-to-video (V2V) translation has enabled realistic resimulation of embodied AI demonstrations, a capability that allows pretrained robot policies to be transferable to new environments without additional data collection. However, prior works can only operate on a single view at a time, while embodied AI tasks are commonly captured from multiple synchronized cameras to support policy learning. Naively applying single-view models independently to each camera leads to inconsistent appearance across views, and standard transformer architectures do not scale to multi-view settings due to the quadratic cost of cross-view attention. We present VideoWeaver, the first multimodal multi-view V2V translation framework. VideoWeaver is initially trained as a single-view flow-based V2V model. To achieve an extension to the multi-view regime, we propose to ground all views in a shared 4D latent space derived from a feed-forward spatial foundation model, namely, Pi3. This encourages view-consistent appearance even under wide baselines and dynamic camera motion. To scale beyond a fixed number of cameras, we train views at distinct diffusion timesteps, enabling the model to learn both joint and conditional view distributions. This in turn allows autoregressive synthesis of new viewpoints conditioned on existing ones. Experiments show superior or similar performance to the state-of-the-art on the single-view translation benchmarks and, for the first time, physically and stylistically consistent multi-view translations, including challenging egocentric and heterogeneous-camera setups central to world randomization for robot learning.

视频生成多视角机器人学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。