用多模态信号控制3D人脸动画,让表情更自然生动。
Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
- 通过对比学习统一文本、音频和情绪标签,实现多源控制。
- 在时间一致性下提升表情多样性,情绪相似度提升21.6%。
- 适合需要高表达力的影视/游戏角色动画开发人员。
基于音频的情感化3D人脸动画面临两大挑战:(1) 依赖单一模态控制信号(视频、文本或情绪标签),未能充分利用其互补优势进行综合情绪操控;(2) 采用确定性回归映射,限制了情感表达与非语言行为的随机性,制约了生成动画的表现力。为此,我们提出一种基于扩散模型的可控高表达力3D人脸动画框架。核心创新包括:(1) 基于FLAME的多模态情绪绑定策略,通过对比学习对齐文本、音频与情绪标签,支持多信号源灵活控制;(2) 采用内容感知注意力与情绪引导层的潜空间扩散模型,增强动作多样性同时保持时序连贯性与自然面部动态。大量实验表明,该方法在多数指标上优于现有方法,情绪相似度提升21.6%,且保持生理上合理的面部运动。
原文摘要 · Abstract (English)
Audio-driven emotional 3D facial animation encounters two significant challenges: (1) reliance on single-modal control signals (videos, text, or emotion labels) without leveraging their complementary strengths for comprehensive emotion manipulation, and (2) deterministic regression-based mapping that constrains the stochastic nature of emotional expressions and non-verbal behaviors, limiting the expressiveness of synthesized animations. To address these challenges, we present a diffusion-based framework for controllable expressive 3D facial animation. Our approach introduces two key innovations: (1) a FLAME-centered multimodal emotion binding strategy that aligns diverse modalities (text, audio, and emotion labels) through contrastive learning, enabling flexible emotion control from multiple signal sources, and (2) an attention-based latent diffusion model with content-aware attention and emotion-guided layers, which enriches motion diversity while maintaining temporal coherence and natural facial dynamics. Extensive experiments demonstrate that our method outperforms existing approaches across most metrics, achieving a 21.6\% improvement in emotion similarity while preserving physiologically plausible facial dynamics. Project Page: https://kangweiiliu.github.io/Control_3D_Animation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。