用4D高斯显式建模动态景观,实现单图生成360度可动视频。
Optimizing 4D Gaussians for Dynamic Scene Video from Single Landscape Images

- 从单张风景图生成多视角图像,通过一致运动估计优化4D高斯
- 在多个数据集上实现真实感动态效果,支持自由视角观看
- 首次将完整3D空间动画化,适合虚拟旅游与数字内容创作
为在风景图中实现真实沉浸感,水体和云彩需在图像内移动并从不同相机视角揭示新场景。近年来兴起的动态场景视频技术结合了单图动画与3D摄影,采用伪3D空间表示,以分层深度图(LDI)隐式建模。LDI将图像按深度分层,使水、云等元素可运动并呈现多视角。然而,连续性元素如流体在离散分层下会削弱深度感知,且可能因相机移动产生畸变。此外,隐式3D建模限制输出仅在2D域,降低灵活性。本文提出从单张图像显式建模完整3D空间,使用4D高斯表示动态场景。框架通过生成多视角图像,建立3D运动以优化4D高斯。核心是统一3D运动估计,通过多视角一致性提升实际运动逼真度。据我们所知,这是首次在单图基础上实现完整3D空间动画化。实验表明,模型在多种风景图上均能生成真实感动态视频,结果详见 https://cvsp-lab.github.io/ICLR2025_3D-MOM/。
原文摘要 · Abstract (English)
To achieve realistic immersion in landscape images, fluids such as water and clouds need to move within the image while revealing new scenes from various camera perspectives. Recently, a field called dynamic scene video has emerged, which combines single image animation with 3D photography. These methods use pseudo 3D space, implicitly represented with Layered Depth Images (LDIs). LDIs separate a single image into depth-based layers, which enables elements like water and clouds to move within the image while revealing new scenes from different camera perspectives. However, as landscapes typically consist of continuous elements, including fluids, the representation of a 3D space separates a landscape image into discrete layers, and it can lead to diminished depth perception and potential distortions depending on camera movement. Furthermore, due to its implicit modeling of 3D space, the output may be limited to videos in the 2D domain, potentially reducing their versatility. In this paper, we propose representing a complete 3D space for dynamic scene video by modeling explicit representations, specifically 4D Gaussians, from a single image. The framework is focused on optimizing 3D Gaussians by generating multi-view images from a single image and creating 3D motion to optimize 4D Gaussians. The most important part of proposed framework is consistent 3D motion estimation, which estimates common motion among multi-view images to bring the motion in 3D space closer to actual motions. As far as we know, this is the first attempt that considers animation while representing a complete 3D space from a single landscape image. Our model demonstrates the ability to provide realistic immersion in various landscape images through diverse experiments and metrics. Extensive experimental results are https://cvsp-lab.github.io/ICLR2025_3D-MOM/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。