arXiv:2608.15110cs.CVcs.AI2026-08

让3D虚拟人脸情绪表达更自然,支持连续情感控制。

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

论文配图:CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation
图 1 · 摘自论文原文
  • 用连续情绪坐标控制面部动画,替代离散情绪分类。
  • 在多个数据集上实现更高唇形同步精度和情感表现力。
  • 适合需要细腻情绪变化的虚拟人、影视特效等场景。

情感驱动的3D说话头生成旨在合成具有准确口型同步的生动面部动画。然而,现有方法多依赖离散情绪类别,难以捕捉情感的连续演变过程,且忽视了语音发音与情感表达之间的时序频率差异。本文提出CETalk,一种基于连续效价-唤醒度(VA)表示的音频驱动3D面部动画框架,实现细粒度情感控制。CETalk通过三个核心组件预测一系列FLAME参数:动态情绪调制模块利用音频线索自适应调节情感强度;多尺度时序建模机制采用并行分支解耦高频发音动作与低频情感动态;动态融合机制通过自适应门控网络整合多尺度特征。为支持训练与评估,我们构建了大规模数据集3D-VA-MEAD,包含自动估算的VA标注与重建的3D面部运动。大量实验证明,CETalk在唇形同步精度与情感表现力方面均优于当前最优方法,同时支持平滑可控的情感过渡。

原文摘要 · Abstract (English)

Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.

3D说话头情感控制音频驱动连续情绪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。