实现人脸表情精准控制的同时保持高保真身份特征。
EmojiDiff: Advanced Facial Expression Control with High Identity Preservation in Portrait Generation
- 采用解耦训练与精调的两阶段框架,分离表情与身份特征。
- 在多种扩散模型上实现高精度表情控制且身份保持率显著提升。
- 适合需要精细表情编辑的数字人、虚拟偶像等应用。
本文旨在实现人物肖像生成中细粒度表情控制与高保真身份保持的统一。由于表情与身份存在相互干扰:(i) 精细的表情控制信号会引入外观语义(如面部轮廓、比例),影响身份一致性;(ii) 即使是粗粒度的表情控制也会作用于面部,导致身份变化。现有方法多依赖粗略控制信号或两阶段推理,未能有效解决此问题。为此,我们提出EmojiDiff,首个端到端实现像素级(RGB级)精细表情控制与高保真身份保持的方法。通过创新的无身份数据迭代(IDI)策略,将身份保持与表情变换过程解耦,生成高质量跨身份表情对,从而在训练中有效分离表达模板中的细粒度表情特征与其他无关信息(如身份、肤色)。随后引入增强身份对比对齐(ICA)进行微调,实现身份与表情信息的快速重建与联合监督,对齐有无表情控制图像的身份表示。实验表明,该方法显著优于现有方案,在多种扩散模型上均表现出优异的表达控制精度与身份保持能力,并具备良好泛化性。
原文摘要 · Abstract (English)
This paper aims to bring fine-grained expression control while maintaining high-fidelity identity in portrait generation. This is challenging due to the mutual interference between expression and identity: (i) fine expression control signals inevitably introduce appearance-related semantics (e.g., facial contours, and ratio), which impact the identity of the generated portrait; (ii) even coarse-grained expression control can cause facial changes that compromise identity, since they all act on the face. These limitations remain unaddressed by previous generation methods, which primarily rely on coarse control signals or two-stage inference that integrates portrait animation. Here, we introduce EmojiDiff, the first end-to-end solution that enables simultaneous control of extremely detailed expression (RGB-level) and high-fidelity identity in portrait generation. To address the above challenges, EmojiDiff adopts a two-stage scheme involving decoupled training and fine-tuning. For decoupled training, we innovate ID-irrelevant Data Iteration (IDI) to synthesize cross-identity expression pairs by dividing and optimizing the processes of maintaining expression and altering identity, thereby ensuring stable and high-quality data generation. Training the model with this data, we effectively disentangle fine expression features in the expression template from other extraneous information (e.g., identity, skin). Subsequently, we present ID-enhanced Contrast Alignment (ICA) for further fine-tuning. ICA achieves rapid reconstruction and joint supervision of identity and expression information, thus aligning identity representations of images with and without expression control. Experimental results demonstrate that our method remarkably outperforms counterparts, achieves precise expression control with highly maintained identity, and generalizes well to various diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。