arXiv:2506.17705cs.CV2025-06被引 3

用视频扩散模型生成随相机移动和物体动态变化的长期视频。

DreamJourney: Perpetual View Generation with Video Diffusion Models

  • 分两阶段:先重建3D场景并生成连贯视图,再注入物体运动。
  • 在10秒长视频上实现视觉一致性与动态细节,优于现有方法。
  • 适合做动态场景模拟、虚拟拍摄或元宇宙内容生成的研究者。

持续视图生成旨在仅基于单张输入图像,沿任意相机轨迹生成长时间视频。现有方法通常使用预训练文本到图像扩散模型合成未见区域内容,但其缺乏3D感知能力,易产生畸变伪影,且仅限于静态3D场景,无法捕捉动态4D世界中的物体运动。为此,我们提出DreamJourney,一个两阶段框架,利用视频扩散模型的世界模拟能力,开启兼具相机运动与物体动态的持续场景视图生成任务。第一阶段:将输入图像提升为3D点云,沿特定相机轨迹渲染部分图像序列;使用视频扩散模型作为生成先验,补全缺失区域并增强序列间视觉连贯性,生成符合3D场景与相机轨迹的跨视角一致视频。同时引入早停与视图填充两种简单有效策略,进一步稳定生成过程并提升画质。第二阶段:利用多模态大语言模型生成描述当前视图中物体运动的文本提示,并通过视频扩散模型实现视图动画化。阶段一与二循环执行,实现持续动态场景视图生成。大量实验表明,DreamJourney在定量与定性上均优于现有最先进方法。

原文摘要 · Abstract (English)

Perpetual view generation aims to synthesize a long-term video corresponding to an arbitrary camera trajectory solely from a single input image. Recent methods commonly utilize a pre-trained text-to-image diffusion model to synthesize new content of previously unseen regions along camera movement. However, the underlying 2D diffusion model lacks 3D awareness and results in distorted artifacts. Moreover, they are limited to generating views of static 3D scenes, neglecting to capture object movements within the dynamic 4D world. To alleviate these issues, we present DreamJourney, a two-stage framework that leverages the world simulation capacity of video diffusion models to trigger a new perpetual scene view generation task with both camera movements and object dynamics. Specifically, in stage I, DreamJourney first lifts the input image to 3D point cloud and renders a sequence of partial images from a specific camera trajectory. A video diffusion model is then utilized as generative prior to complete the missing regions and enhance visual coherence across the sequence, producing a cross-view consistent video adheres to the 3D scene and camera trajectory. Meanwhile, we introduce two simple yet effective strategies (early stopping and view padding) to further stabilize the generation process and improve visual quality. Next, in stage II, DreamJourney leverages a multimodal large language model to produce a text prompt describing object movements in current view, and uses video diffusion model to animate current view with object movements. Stage I and II are repeated recurrently, enabling perpetual dynamic scene view generation. Extensive experiments demonstrate the superiority of our DreamJourney over state-of-the-art methods both quantitatively and qualitatively. Our project page: https://dream-journey.vercel.app.

视频生成扩散模型动态场景3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。