仅用一张图生成动态4D内容,提升画面连贯性与真实感。
MVG4D: Image Matrix-Based Multi-View and Motion Generation for 4D Content Creation from a Single Image
- 通过图像矩阵生成多视角图像,指导3D点云重建。
- 在Objaverse上实现更高清晰度、更低闪烁,速度优于现有方法。
- 适合需要高效可控4D生成的AR/VR开发人员使用。
生成模型的发展已推动数字内容从2D图像扩展至复杂的3D和4D场景。尽管进展显著,高质量且时序一致的动态4D内容生成仍是挑战。本文提出MVG4D框架,通过结合多视角合成与4D高斯溅射(4D GS),从单张静态图像生成动态4D内容。核心在于图像矩阵模块,可生成时序一致且空间多样的多视角图像,为下游3D和4D重建提供丰富监督信号。这些多视角图像用于优化3D高斯点云,并通过轻量形变网络扩展至时间维度。该方法有效提升了时序一致性、几何保真度与视觉真实感,缓解了先前4D GS方法中存在的运动不连续和背景退化问题。在Objaverse数据集上的大量实验表明,MVG4D在CLIP-I、PSNR、FVD指标及时间效率上均超越主流基线。显著减少闪烁伪影,增强各视角与时间维度的结构细节,为更沉浸的AR/VR体验提供支持。MVG4D为从极少输入中高效可控生成4D内容开辟了新方向。
原文摘要 · Abstract (English)
Advances in generative modeling have significantly enhanced digital content creation, extending from 2D images to complex 3D and 4D scenes. Despite substantial progress, producing high-fidelity and temporally consistent dynamic 4D content remains a challenge. In this paper, we propose MVG4D, a novel framework that generates dynamic 4D content from a single still image by combining multi-view synthesis with 4D Gaussian Splatting (4D GS). At its core, MVG4D employs an image matrix module that synthesizes temporally coherent and spatially diverse multi-view images, providing rich supervisory signals for downstream 3D and 4D reconstruction. These multi-view images are used to optimize a 3D Gaussian point cloud, which is further extended into the temporal domain via a lightweight deformation network. Our method effectively enhances temporal consistency, geometric fidelity, and visual realism, addressing key challenges in motion discontinuity and background degradation that affect prior 4D GS-based methods. Extensive experiments on the Objaverse dataset demonstrate that MVG4D outperforms state-of-the-art baselines in CLIP-I, PSNR, FVD, and time efficiency. Notably, it reduces flickering artifacts and sharpens structural details across views and time, enabling more immersive AR/VR experiences. MVG4D sets a new direction for efficient and controllable 4D generation from minimal inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。