用单图生成全景视频并转为4D场景,提升VR/AR沉浸感
HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
- 构建360度全景视频数据集,训练两阶段扩散模型生成高质量全景视频
- 通过时空深度估计将视频转为4D点云,重建时空一致的4D场景
- 适合做虚拟现实、增强现实的4D内容生成,尤其关注全景动态场景
扩散模型的快速发展为虚拟现实与增强现实技术带来变革机遇,这些应用通常需要场景级4D资产以实现沉浸式体验。然而现有扩散模型多聚焦静态3D场景或物体级动态,难以支持真正沉浸的应用。为此,我们提出HoloTime框架,利用视频扩散模型从单个提示或参考图像生成全景视频,并结合360度4D场景重建方法,将生成的全景视频无缝转化为4D资产,实现用户全沉浸式4D体验。具体而言,为提升视频扩散模型生成高保真全景视频的能力,我们构建了首个适用于下游4D场景重建任务的360World全景视频数据集。基于此数据集,提出Panoramic Animator——一个两阶段图像到视频的扩散模型,可将全景图像转化为高质量全景视频。随后,提出Panoramic Space-Time Reconstruction,利用时空深度估计方法将生成的全景视频转换为4D点云,并优化4D高斯泼溅表示,实现空间与时间上一致的4D场景重建。通过与现有方法对比分析,验证了该方法在全景视频生成与4D场景重建上的优越性,证明其能创造更具吸引力与真实感的沉浸环境,显著提升VR/AR应用中的用户体验。
原文摘要 · Abstract (English)
The rapid advancement of diffusion models holds the promise of revolutionizing the application of VR and AR technologies, which typically require scene-level 4D assets for user experience. Nonetheless, existing diffusion models predominantly concentrate on modeling static 3D scenes or object-level dynamics, constraining their capacity to provide truly immersive experiences. To address this issue, we propose HoloTime, a framework that integrates video diffusion models to generate panoramic videos from a single prompt or reference image, along with a 360-degree 4D scene reconstruction method that seamlessly transforms the generated panoramic video into 4D assets, enabling a fully immersive 4D experience for users. Specifically, to tame video diffusion models for generating high-fidelity panoramic videos, we introduce the 360World dataset, the first comprehensive collection of panoramic videos suitable for downstream 4D scene reconstruction tasks. With this curated dataset, we propose Panoramic Animator, a two-stage image-to-video diffusion model that can convert panoramic images into high-quality panoramic videos. Following this, we present Panoramic Space-Time Reconstruction, which leverages a space-time depth estimation method to transform the generated panoramic videos into 4D point clouds, enabling the optimization of a holistic 4D Gaussian Splatting representation to reconstruct spatially and temporally consistent 4D scenes. To validate the efficacy of our method, we conducted a comparative analysis with existing approaches, revealing its superiority in both panoramic video generation and 4D scene reconstruction. This demonstrates our method's capability to create more engaging and realistic immersive environments, thereby enhancing user experiences in VR and AR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。