仅用一张图就能精准保留人脸身份,生成逼真表情姿态变化图像。
InstaFace: Identity-Preserving Facial Editing with Single Image Inference
- 基于扩散模型,引入无参数3D视角引导网络,实现单图身份保持。
- 在身份保真度上超越多个前沿方法,面部特征更自然真实。
- 适合数字人、AR/VR、个性化内容创作等需要高保真人脸编辑的场景。
人脸外观编辑对数字虚拟形象、AR/VR及个性化内容创作至关重要,但生成模型在有限数据下难以保持身份一致性。传统方法需多张图像,仍存在脸型扭曲、发型错位或过度平滑等问题。为此,我们提出基于扩散模型的InstaFace框架,仅需单张图像即可生成逼真图像并保持身份特征。核心创新在于引入无需额外训练参数的3DMM条件融合引导网络,结合人脸识别模型与预训练视觉语言模型的特征嵌入,有效保留身份、背景、发丝及配饰等上下文信息。定量评估显示,本方法在身份保真度、照片级真实感以及姿态、表情、光照控制方面均优于多个先进方法。
原文摘要 · Abstract (English)
Facial appearance editing is crucial for digital avatars, AR/VR, and personalized content creation, driving realistic user experiences. However, preserving identity with generative models is challenging, especially in scenarios with limited data availability. Traditional methods often require multiple images and still struggle with unnatural face shifts, inconsistent hair alignment, or excessive smoothing effects. To overcome these challenges, we introduce a novel diffusion-based framework, InstaFace, to generate realistic images while preserving identity using only a single image. Central to InstaFace, we introduce an efficient guidance network that harnesses 3D perspectives by integrating multiple 3DMM-based conditionals without introducing additional trainable parameters. Moreover, to ensure maximum identity retention as well as preservation of background, hair, and other contextual features like accessories, we introduce a novel module that utilizes feature embeddings from a facial recognition model and a pre-trained vision-language model. Quantitative evaluations demonstrate that our method outperforms several state-of-the-art approaches in terms of identity preservation, photorealism, and effective control of pose, expression, and lighting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。