arXiv:2508.20615cs.CV2025-08被引 3

用文字精准控制表情,生成自然同步的动态肖像视频。

EmoCAST: Emotional Talking Portrait via Emotive Text Description

  • 通过情感文本注意力模块实现文本驱动的表情控制。
  • 在真实场景数据上训练,实现高保真口型同步与细腻表情。
  • 适合需要情感化虚拟人交互的开发者和内容创作者。

情感化说话头合成旨在生成具有生动表情的动态肖像视频。现有方法在控制灵活性、动作自然度和表情质量方面仍存在局限,且当前数据集多为实验室采集,进一步限制了实际应用。为此,我们提出 EmoCAST,一种基于扩散模型的文本驱动情感化说话头框架。其贡献包括:(1)支持有效文本控制的架构模块;(2)一个大规模野外场景的情感说话头数据集;(3)提升性能的训练策略。具体而言,在外观建模中,通过文本引导的情感注意力模块整合情绪提示,增强空间知识以提升情感理解。为强化音频-情绪对齐,引入情感音频注意力模块,捕捉可控情绪与驱动音频之间的相互作用,生成情绪感知特征以指导精确的面部动作合成。此外,构建了一个包含情感文本描述的大规模野外情感说话头数据集,并据此提出情绪感知采样策略和渐进式功能训练策略,显著提升模型对细微表达特征的捕捉能力与口型同步精度。整体上,EmoCAST 在生成逼真、富有情感且音频同步的说话头视频方面达到领先水平。

原文摘要 · Abstract (English)

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available datasets are mainly collected in lab settings, further exacerbating these shortcomings and hindering real-world deployment. To address these challenges, we propose EmoCAST, a diffusion-based talking head framework for precise, text-driven emotional synthesis. Its contributions are threefold: (1) architectural modules that enable effective text control; (2) an emotional talking-head dataset that expands the framework's ability; and (3) training strategies that further improve performance. Specifically, for appearance modeling, emotional prompts are integrated through a text-guided emotive attention module, enhancing spatial knowledge to improve emotion understanding. To strengthen audio-emotion alignment, we introduce an emotive audio attention module to capture the interplay between controlled emotion and driving audio, generating emotion-aware features to guide precise facial motion synthesis. Additionally, we construct a large-scale, in-the-wild emotional talking head dataset with emotive text descriptions to optimize the framework's performance. Based on this dataset, we propose an emotion-aware sampling strategy and a progressive functional training strategy that improve the model's ability to capture nuanced expressive features and achieve accurate lip-sync. Overall, EmoCAST achieves state-of-the-art performance in generating realistic, emotionally expressive, and audio-synchronized talking-head videos. Project Page: https://github.com/GVCLab/EmoCAST

情感生成说话头扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。