用单目视频生成高保真4D内容,时空一致且支持文本/图像引导。
Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
- 融合多视角渲染与视频扩散模型,实现时空一致性优化。
- 在多个公开数据集上达到当前最佳性能,细节保留更完整。
- 适合数字人、AR/VR等需要可控4D生成的应用场景。
从单目视频生成高质量4D内容(如数字人、AR/VR应用)面临时空一致性难保障、细节丢失及用户指令融合困难等问题。为此,我们提出Splat4D框架,通过多视角渲染、不一致性检测、视频扩散模型与非对称U-Net协同优化,实现高保真4D内容生成。在多个公开基准测试中,Splat4D持续表现出最优性能,验证了方法的有效性。此外,其多功能性在文本/图像条件4D生成、4D人体生成及文本引导编辑任务中均取得连贯结果,准确遵循用户指令。
原文摘要 · Abstract (English)
Generating high-quality 4D content from monocular videos for applications such as digital humans and AR/VR poses challenges in ensuring temporal and spatial consistency, preserving intricate details, and incorporating user guidance effectively. To overcome these challenges, we introduce Splat4D, a novel framework enabling high-fidelity 4D content generation from a monocular video. Splat4D achieves superior performance while maintaining faithful spatial-temporal coherence by leveraging multi-view rendering, inconsistency identification, a video diffusion model, and an asymmetric U-Net for refinement. Through extensive evaluations on public benchmarks, Splat4D consistently demonstrates state-of-the-art performance across various metrics, underscoring the efficacy of our approach. Additionally, the versatility of Splat4D is validated in various applications such as text/image conditioned 4D generation, 4D human generation, and text-guided content editing, producing coherent outcomes following user instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。