仅用一张人脸图生成可实时驱动的高保真3D虚拟人像
FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image

- 结合视觉基础模型与扩散模型,从单张图像推断完整3D头部结构
- 无需个性化优化,生成结果在身份保留和细节质量上超越现有方法
- 基于3D高斯表示,实现真实感渲染与实时表情驱动
我们提出FiCA,一种前向、即时的3D高斯编码虚拟人像生成流程,仅需单张肖像图即可生成逼真的可驱动虚拟人像。由于单张图像提供的视觉信息有限,准确重建人类头部的3D外观与几何结构极具挑战。为此,我们设计了一种新系统,融合以人为中心的视觉基础模型与扩散模型,充分利用部分视觉观测生成高保真虚拟人像。所提出的扩散模型学习从这些不完整观测到完整且真实的3D网格重建的生成映射。此外,我们引入前向网格精修网络,提升生成人像的保真度与身份一致性,无需个性化的测试时优化。通过利用通用先验模型将生成网格解码为一组3D高斯,我们生成可实时驱动的逼真3D高斯虚拟人像。实验表明,该前向方法生成的虚拟人像忠实还原多样身份,视觉质量优于近期竞争方法。
原文摘要 · Abstract (English)
We introduce FiCA, a Feed-forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Generating a photorealistic and drivable avatar from just a single image is significantly challenging due to the limited visual information available to accurately infer the 3D appearance and geometry of human heads. To address this, we develop a novel system that combines human-centric vision foundation models with a diffusion model. This system is designed to fully exploit partial visual observations to generate lifelike human avatars. Our proposed diffusion model learns a generative mapping from these partial observations to complete and authentic 3D mesh reconstruction. Additionally, we introduce a feed-forward mesh refinement network that enhances the fidelity and identity preservation of the generated avatars, eliminating the need for person-specific test-time optimization. By leveraging a universal prior model that decodes a generated mesh into a set of 3D Gaussians, we generate a photorealistic 3D Gaussian avatar, capable of being driven with novel expressions in real-time. Our experiments demonstrate that the avatars generated by our feed-forward approach faithfully represent diverse identities and surpass the visual quality of avatars produced by recent competing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。