单图生成可动画3D人像,支持多种风格与细节保留
SOAP: Style-Omniscient Animatable Portraits
- 用多视角扩散模型训练24000个不同风格头像,生成带骨骼的3D avatar
- 通过可微渲染优化FLAME网格,保持拓扑和绑定结构,支持面部动作控制
- 能还原发辫、配饰等复杂细节,适合影视/游戏角色快速建模
从单张图像生成可动画的3D虚拟人像仍面临风格限制(写实、卡通、动漫)及对配饰、发型处理困难的问题。尽管3D扩散模型在通用物体单视图重建上取得进展,但输出常缺乏动画控制或出现伪影,源于领域差异。本文提出SOAP框架,可从任意人像生成带骨骼、拓扑一致的3D虚拟人像。方法基于在24,000个具有多种风格的3D头部数据上训练的多视角扩散模型,并结合自适应优化流程,在保持拓扑和绑定的前提下,通过可微渲染变形FLAME网格。生成的纹理化人像支持基于FACS的动画,集成眼球与牙齿,可精确保留辫发、配饰等细节。大量实验表明,该方法在单视图头部建模及基于扩散模型的图像到3D生成方面均优于当前最优技术。代码与数据已公开于https://github.com/TingtingLiao/soap。
原文摘要 · Abstract (English)
Creating animatable 3D avatars from a single image remains challenging due to style limitations (realistic, cartoon, anime) and difficulties in handling accessories or hairstyles. While 3D diffusion models advance single-view reconstruction for general objects, outputs often lack animation controls or suffer from artifacts because of the domain gap. We propose SOAP, a style-omniscient framework to generate rigged, topology-consistent avatars from any portrait. Our method leverages a multiview diffusion model trained on 24K 3D heads with multiple styles and an adaptive optimization pipeline to deform the FLAME mesh while maintaining topology and rigging via differentiable rendering. The resulting textured avatars support FACS-based animation, integrate with eyeballs and teeth, and preserve details like braided hair or accessories. Extensive experiments demonstrate the superiority of our method over state-of-the-art techniques for both single-view head modeling and diffusion-based generation of Image-to-3D. Our code and data are publicly available for research purposes at https://github.com/TingtingLiao/soap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。