arXiv:2503.19906cs.CV2025-03CVPR被引 16

用一张任意风格的头像图,生成可动的3D虚拟人像。

AvatarArtist: Open-Domain 4D Avatarization

  • 用参数化三平面作中间表示,结合GAN与扩散模型训练。
  • 在多种风格图像上都能生成高质量4D avatar,鲁棒性强。
  • 适合做虚拟形象、数字人、游戏角色等领域的开发者使用。

本研究聚焦开放域4D虚拟人像生成,旨在从任意风格的头像图像中生成4D虚拟人像。采用参数化三平面作为中间4D表示,并提出一种结合生成对抗网络(GAN)与扩散模型的实用训练范式。观察发现,4D GAN虽无需监督即可在图像与三平面间映射,但难以处理多样数据分布;为此引入稳健的2D扩散模型先验,辅助GAN跨域迁移能力。二者协同构建多领域图像-三平面数据集,推动通用4D虚拟人像生成器的发展。大量实验表明,所提模型AvatarArtist能生成高质量4D虚拟人像,对多种源图像领域具有强鲁棒性。代码、数据与模型将公开,以促进后续研究。

原文摘要 · Abstract (English)

This work focuses on open-domain 4D avatarization, with the purpose of creating a 4D avatar from a portrait image in an arbitrary style. We select parametric triplanes as the intermediate 4D representation and propose a practical training paradigm that takes advantage of both generative adversarial networks (GANs) and diffusion models. Our design stems from the observation that 4D GANs excel at bridging images and triplanes without supervision yet usually face challenges in handling diverse data distributions. A robust 2D diffusion prior emerges as the solution, assisting the GAN in transferring its expertise across various domains. The synergy between these experts permits the construction of a multi-domain image-triplane dataset, which drives the development of a general 4D avatar creator. Extensive experiments suggest that our model, AvatarArtist, is capable of producing high-quality 4D avatars with strong robustness to various source image domains. The code, the data, and the models will be made publicly available to facilitate future studies.

虚拟人像4D生成扩散模型GAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。