用视频生成数据训练3D高保真角色,单图也能动起来
SVAD: From Single Image to 3D Avatar via Synthetic Data Generation with Video Diffusion and Data Augmentation
- 用扩散模型生成带身份保持的合成视频数据
- 在单张图像下实现高保真细节与视角一致性
- 适合需要高效高质量角色生成的研究者
从单张图像生成可动画化高保真3D人体角色仍是计算机视觉中的难题,因单一视角难以重建完整3D信息。现有方法存在明显局限:3D高斯泼溅(3DGS)需多视角或视频序列才能获得高质量结果,而视频扩散模型虽可单图生成动画,却难保持身份一致性和细节。本文提出SVAD,通过视频扩散生成合成训练数据,结合身份保持与图像修复模块增强数据质量,并利用优化后数据训练3DGS角色。实验表明,SVAD在身份一致性与新姿态/视角下的细节保留上优于当前最佳单图方法,且支持实时渲染。该方法突破了传统3DGS对密集单目或多视图数据的依赖。定量与定性对比显示,本方法在多个指标上显著超越基线模型。通过融合扩散模型的生成能力与3DGS的高质量输出和高效渲染优势,建立了一种基于单图生成高保真角色的新范式。
原文摘要 · Abstract (English)
Creating high-quality animatable 3D human avatars from a single image remains a significant challenge in computer vision due to the inherent difficulty of reconstructing complete 3D information from a single viewpoint. Current approaches face a clear limitation: 3D Gaussian Splatting (3DGS) methods produce high-quality results but require multiple views or video sequences, while video diffusion models can generate animations from single images but struggle with consistency and identity preservation. We present SVAD, a novel approach that addresses these limitations by leveraging complementary strengths of existing techniques. Our method generates synthetic training data through video diffusion, enhances it with identity preservation and image restoration modules, and utilizes this refined data to train 3DGS avatars. Comprehensive evaluations demonstrate that SVAD outperforms state-of-the-art (SOTA) single-image methods in maintaining identity consistency and fine details across novel poses and viewpoints, while enabling real-time rendering capabilities. Through our data augmentation pipeline, we overcome the dependency on dense monocular or multi-view training data typically required by traditional 3DGS approaches. Extensive quantitative, qualitative comparisons show our method achieves superior performance across multiple metrics against baseline models. By effectively combining the generative power of diffusion models with both the high-quality results and rendering efficiency of 3DGS, our work establishes a new approach for high-fidelity avatar generation from a single image input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。