arXiv:2507.12761cs.CVcs.AI2025-07

让AI先理解情绪再绘脸,生成更自然的动态表情视频。

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

  • 用思维链解析情绪,将抽象标签转为肌肉动作描述。
  • 通过渐进式去噪优化微表情,实现精细表情控制。
  • 适合需要高情感真实性的交互系统开发者。

情感驱动的说话头生成是计算机视觉与多模态AI交叉领域的关键研究方向,其核心价值在于通过沉浸式、共情化交互提升人机互动体验。随着多模态大模型的发展,驱动信号从音视频转向更灵活的文本。然而现有文本驱动方法依赖预定义离散情绪标签,过度简化真实面部肌肉运动的动态复杂性,难以实现自然的情感表现力。本文提出Think-Before-Draw框架,解决两大挑战:(1) 情绪语义深度解析——创新引入思维链(Chain-of-Thought, CoT),将抽象情绪标签转化为生理学基础的面部肌肉运动描述,实现从高层语义到可执行运动特征的映射;(2) 细粒度表现力优化——借鉴艺术家肖像绘画过程,提出渐进式引导去噪策略,采用“全局情绪定位—局部肌肉控制”机制,精细化调控生成视频中的微表情动态。实验表明,该方法在主流基准MEAD和HDTF上达到领先性能。此外,我们收集了一组肖像图像用于评估模型的零样本生成能力。

原文摘要 · Abstract (English)

Emotional talking-head generation has emerged as a pivotal research area at the intersection of computer vision and multimodal artificial intelligence, with its core value lying in enhancing human-computer interaction through immersive and empathetic engagement.With the advancement of multimodal large language models, the driving signals for emotional talking-head generation has shifted from audio and video to more flexible text. However, current text-driven methods rely on predefined discrete emotion label texts, oversimplifying the dynamic complexity of real facial muscle movements and thus failing to achieve natural emotional expressiveness.This study proposes the Think-Before-Draw framework to address two key challenges: (1) In-depth semantic parsing of emotions--by innovatively introducing Chain-of-Thought (CoT), abstract emotion labels are transformed into physiologically grounded facial muscle movement descriptions, enabling the mapping from high-level semantics to actionable motion features; and (2) Fine-grained expressiveness optimization--inspired by artists' portrait painting process, a progressive guidance denoising strategy is proposed, employing a "global emotion localization--local muscle control" mechanism to refine micro-expression dynamics in generated videos.Our experiments demonstrate that our approach achieves state-of-the-art performance on widely-used benchmarks, including MEAD and HDTF. Additionally, we collected a set of portrait images to evaluate our model's zero-shot generation capability.

情感生成表达控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。