让虚拟人脸随文本情绪变化自然流动,生成更真实的情感表达视频。
Text-Driven Emotionally Continuous Talking Face Generation
- 通过时间敏感的情绪波动建模,动态控制表情变化序列。
- 实现平滑的情绪过渡,保持高画质与动作真实性。
- 适合需要情感连续性的数字人、影视动画等场景。
说话人脸生成(TFG)致力于创建真实且富有情感表达的数字人脸。尽管以往方法已能生成自然的面部动作,但通常仅表达单一固定情绪,难以模拟人类在传递信息时持续变化的自然表情。为此,我们提出新任务——情感连续说话人脸生成(EC-TFG),以文本片段和包含多变情绪描述作为驱动数据,目标是生成一个人在讲述该文本的同时,其表情随情绪描述自然演变的视频。为此,我们设计了定制化模型TIE-TFG,创新性地采用时间敏感情绪波动建模机制,能够根据输入文本生成对应的情绪变化序列,从而驱动合成视频中连续的表情演化。大量实验表明,本方法在多种情绪状态下均能实现平滑的情绪过渡,并保持高质量视觉效果与动作真实感。
原文摘要 · Abstract (English)
Talking Face Generation (TFG) strives to create realistic and emotionally expressive digital faces. While previous TFG works have mastered the creation of naturalistic facial movements, they typically express a fixed target emotion in synthetic videos and lack the ability to exhibit continuously changing and natural expressions like humans do when conveying information. To synthesize realistic videos, we propose a novel task called Emotionally Continuous Talking Face Generation (EC-TFG), which takes a text segment and an emotion description with varying emotions as driving data, aiming to generate a video where the person speaks the text while reflecting the emotional changes within the description. Alongside this, we introduce a customized model, i.e., Temporal-Intensive Emotion Modulated Talking Face Generation (TIE-TFG), which innovatively manages dynamic emotional variations by employing Temporal-Intensive Emotion Fluctuation Modeling, allowing it to provide emotion variation sequences corresponding to the input text to drive continuous facial expression changes in synthesized videos. Extensive evaluations demonstrate our method's exceptional ability to produce smooth emotion transitions and uphold high-quality visuals and motion authenticity across diverse emotional states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。