arXiv:2605.29316cs.CV2026-05

让3D人脸动画随文本描述自由切换风格和情绪。

CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation

论文配图:CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation
图 1 · 摘自论文原文
  • 输入语音+文本风格/情绪描述,实现风格与情绪分离控制
  • 支持推理时动态调整情绪,适配语音情感变化
  • 构建大规模图文标注数据集,提升风格化表现力

基于音频的3D人脸动画旨在从任意音频片段生成同步的口型动作和生动的表情。现有方法虽能生成同步口型,但通常依赖预设的身份或风格隐向量,限制了用户对说话风格的自由控制。此外,对整个音频段使用固定风格或身份,导致面部动画风格无法随音频情感内容变化。为此,我们重新审视风格与情感之间的纠缠关系,构建了一个包含风格与情感文本描述的大规模数据集,并提出一种新型对话头生成框架,可分别控制风格与情感。模型输入包括说话风格和角色情绪的文本描述以及驱动音频流,实现实时生成高度同步的口型动作和匹配描述的表情。此外,模型支持推理过程中的动态情绪控制,能够处理语音中情感发生变化的场景。

原文摘要 · Abstract (English)

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or style latent features, which limits users' ability to freely control speaking styles. Moreover, applying a fixed style or identity to an entire audio segment typically results in facial animation styles that do not adapt to the emotional content of the audio. To address these challenges, we revisit the entanglement between style and emotion, construct a large-scale dataset with textual descriptions of both style and emotion, and propose a novel talking head generation framework that enables separate control over style and emotion. Our model takes as input both textual descriptions of speaking style and character emotion, as well as the driving audio stream, enabling real-time generation of highly synchronized lip movements and facial expressions that match the provided descriptions. Furthermore, our model supports dynamic emotion control during inference, allowing it to handle scenarios where the target emotion changes throughout the speech.

3D动画风格控制情感生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。