让人脸生成同时精准控制表情、姿态和情绪,且不破坏身份特征。
FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion
- 在扩散模型中加入轻量控制模块,独立调节面部姿态、表情与情绪。
- 在相同身份下生成多样化人脸,控制精度显著优于现有方法。
- 适合需要精细调控人脸属性的虚拟形象、影视制作等场景。
人脸图像包含丰富的信息,涵盖稳定的身份特征与可变的属性如姿态、表情和情绪。尽管近期图像生成技术已实现高质量的身份条件化人脸合成,但对非身份属性的精确控制仍具挑战性,且实现身份与可变因素的解耦尤为困难。为此,我们提出一种新型的身份条件化扩散模型,引入两个轻量级控制模块,可在不损害身份保留的前提下独立操控面部姿态、表情与情绪。这些模块嵌入基础扩散模型的交叉注意力层中,以极小参数开销实现精准属性控制。此外,我们设计的训练策略通过身份特征与各非身份控制特征间的交叉注意力,促使身份特征与控制信号保持正交,从而增强可控性与生成多样性。定量分析、定性评估及感知用户研究均表明,本方法在姿态、表情与情绪控制精度方面超越现有方法,同时在仅以身份为条件时提升了生成多样性。
原文摘要 · Abstract (English)
Human facial images encode a rich spectrum of information, encompassing both stable identity-related traits and mutable attributes such as pose, expression, and emotion. While recent advances in image generation have enabled high-quality identity-conditional face synthesis, precise control over non-identity attributes remains challenging, and disentangling identity from these mutable factors is particularly difficult. To address these limitations, we propose a novel identity-conditional diffusion model that introduces two lightweight control modules designed to independently manipulate facial pose, expression, and emotion without compromising identity preservation. These modules are embedded within the cross-attention layers of the base diffusion model, enabling precise attribute control with minimal parameter overhead. Furthermore, our tailored training strategy, which leverages cross-attention between the identity feature and each non-identity control feature, encourages identity features to remain orthogonal to control signals, enhancing controllability and diversity. Quantitative and qualitative evaluations, along with perceptual user studies, demonstrate that our method surpasses existing approaches in terms of control accuracy over pose, expression, and emotion, while also improving generative diversity under identity-only conditioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。