用混合形状参数引导扩散模型,精准控制人脸表情且保持身份一致
ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
- 通过表情交叉注意力模块,用FLAME参数实现表情精准控制
- 能生成微表情和流畅过渡,超越基础情绪表达
- 支持真实图像的表情编辑,适合影视动画与虚拟角色生成
面向人工智能叙事的人类中心生成模型需兼顾身份一致性与表演精确控制。现有基于扩散模型的方法虽能保持面部身份,但在不牺牲身份的前提下实现细粒度表情控制仍具挑战。本文提出一种扩散框架,可忠实重现任意主体在特定表情下的形象。基于身份一致的面部基础模型,采用由FLAME混合形状参数引导的组合式设计,引入表达交叉注意力模块实现显式控制。在包含丰富表情变化的图像与视频混合数据集上训练,该适配器可泛化至细微微表情与动态表达转换,弥补先前研究不足。此外,可插拔的参考适配器可在合成过程中从参考帧迁移外观,实现真实图像的表情编辑。大量定量与定性评估表明,本模型在定制化、身份一致的表情生成上优于现有方法。代码与模型见https://github.com/foivospar/Arc2Face。
原文摘要 · Abstract (English)
Human-centric generative models designed for AI-driven storytelling must bring together two core capabilities: identity consistency and precise control over human performance. While recent diffusion-based approaches have made significant progress in maintaining facial identity, achieving fine-grained expression control without compromising identity remains challenging. In this work, we present a diffusion-based framework that faithfully reimagines any subject under any particular facial expression. Building on an ID-consistent face foundation model, we adopt a compositional design featuring an expression cross-attention module guided by FLAME blendshape parameters for explicit control. Trained on a diverse mixture of image and video data rich in expressive variation, our adapter generalizes beyond basic emotions to subtle micro-expressions and expressive transitions, overlooked by prior works. In addition, a pluggable Reference Adapter enables expression editing in real images by transferring the appearance from a reference frame during synthesis. Extensive quantitative and qualitative evaluations show that our model outperforms existing methods in tailored and identity-consistent expression generation. Code and models can be found at https://github.com/foivospar/Arc2Face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。