单图生成高保真3D虚拟人,靠人脸大模型引导
Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance
- 用单张图+人脸大模型做3D头像生成,结合改进的3DGS
- 生成结果保持与模板网格的稠密对应,支持表情控制
- 适合需要快速生成个性化3D角色的创作者和游戏开发者
受3D高斯泼溅(3DGS)在多视角场景重建中的有效性及2D人体基础模型兴起的启发,我们提出Arc2Avatar,首个基于扩散模型、仅需单张图像输入并利用人脸基础模型作为引导的3D虚拟人生成方法。通过在合成数据上微调并修改条件机制,扩展该模型以生成多视角人脸;生成的3D头像与人体面部网格模板保持稠密对应,支持基于混合形状的表情生成。该效果通过改进的3DGS方法、连接性正则化及针对任务设计的初始化实现。此外,我们提出可选的高效基于扩散的校正步骤,进一步提升表情的真实感与多样性。实验表明,Arc2Avatar在真实感与身份保留方面达到当前最优水平,通过极低的引导强度即可解决色彩问题,得益于强大的身份先验与初始化策略,且不损失细节。更多资源请访问 https://arc2avatar.github.io。
原文摘要 · Abstract (English)
Inspired by the effectiveness of 3D Gaussian Splatting (3DGS) in reconstructing detailed 3D scenes within multi-view setups and the emergence of large 2D human foundation models, we introduce Arc2Avatar, the first SDS-based method utilizing a human face foundation model as guidance with just a single image as input. To achieve that, we extend such a model for diverse-view human head generation by fine-tuning on synthetic data and modifying its conditioning. Our avatars maintain a dense correspondence with a human face mesh template, allowing blendshape-based expression generation. This is achieved through a modified 3DGS approach, connectivity regularizers, and a strategic initialization tailored for our task. Additionally, we propose an optional efficient SDS-based correction step to refine the blendshape expressions, enhancing realism and diversity. Experiments demonstrate that Arc2Avatar achieves state-of-the-art realism and identity preservation, effectively addressing color issues by allowing the use of very low guidance, enabled by our strong identity prior and initialization strategy, without compromising detail. Please visit https://arc2avatar.github.io for more resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。