arXiv:2601.05239cs.CV2026-01被引 9

让生成视频在多视角下保持时空一致,支持精准相机控制。

Plenoptic Video Generation

  • 用自回归框架+相机引导检索,让生成视频跨视角同步
  • 在Basic和Agibot上实现最优视图一致性与视觉保真度
  • 适合机器人操控等需多视角连贯生成的场景

相机控制的生成式视频重渲染方法(如 ReCamMaster)已取得显著进展。然而,尽管在单视角设置中表现优异,这些方法在多视角场景中常难以保持一致性。由于生成模型固有的随机性,幻觉区域的时空一致性仍具挑战。为此,我们提出 PlenopticDreamer,一种通过同步生成幻觉以维持时空记忆的框架。核心思想是训练一个视频条件化的多输入单输出自回归模型,并结合相机引导的视频检索策略,从前期生成中动态选择关键视频作为条件输入。此外,训练过程引入渐进式上下文扩展以提升收敛性,自条件机制增强对长程视觉退化的鲁棒性,以及长视频条件机制以支持更长时间的视频生成。在 Basic 与 Agibot 基准上的大量实验表明,PlenopticDreamer 实现了最先进的视频重渲染性能,展现出卓越的视图同步能力、高保真视觉效果、精确的相机控制,以及多样化的视图变换(如机器人操作中的第三人称到第三人称,头视图到夹持器视图)。项目页面:https://research.nvidia.com/labs/dir/plenopticdreamer/

原文摘要 · Abstract (English)

Camera-controlled generative video re-rendering methods, such as ReCamMaster, have achieved remarkable progress. However, despite their success in single-view setting, these works often struggle to maintain consistency across multi-view scenarios. Ensuring spatio-temporal coherence in hallucinated regions remains challenging due to the inherent stochasticity of generative models. To address it, we introduce PlenopticDreamer, a framework that synchronizes generative hallucinations to maintain spatio-temporal memory. The core idea is to train a multi-in-single-out video-conditioned model in an autoregressive manner, aided by a camera-guided video retrieval strategy that adaptively selects salient videos from previous generations as conditional inputs. In addition, Our training incorporates progressive context-scaling to improve convergence, self-conditioning to enhance robustness against long-range visual degradation caused by error accumulation, and a long-video conditioning mechanism to support extended video generation. Extensive experiments on the Basic and Agibot benchmarks demonstrate that PlenopticDreamer achieves state-of-the-art video re-rendering, delivering superior view synchronization, high-fidelity visuals, accurate camera control, and diverse view transformations (e.g., third-person to third-person, and head-view to gripper-view in robotic manipulation). Project page: https://research.nvidia.com/labs/dir/plenopticdreamer/

视频生成多视角相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。