arXiv:2604.13335cs.CV2026-04

用语音情绪分段实现精细表情控制,让3D说话头更自然生动。

SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization

  • 通过帧级语音情绪识别,实时预测情绪类别与强度。
  • 在EmoVOCA数据集上实现低几何误差与平滑表情过渡。
  • 适合需要情感化3D人脸动画的影视、虚拟主播应用。

我们提出SEDTalker,一种基于帧级语音情绪分段的语音驱动3D面部动画框架,实现细粒度的表情控制。与依赖语句级或人工标注情绪标签的方法不同,该方法直接从语音中预测时序密集的情绪类别和强度,支持表情随时间连续调节。分段后的情绪信号被编码为可学习嵌入,并用于条件控制基于混合Transformer-Mamba架构的3D动画模型。该设计有效解耦语言内容与情感风格,同时保持身份一致性和时间连贯性。我们在大规模多语料语音情绪分段数据集及EmoVOCA情感3D面部动画数据集上进行评估,定量结果表明帧级情绪识别性能优异,几何与时间重建误差低;定性结果显示表情过渡自然且表达可控。这些发现证明了帧级情绪分段在生成富有表现力且可调控的3D说话头中的有效性。

原文摘要 · Abstract (English)

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely on utterance-level or manually specified emotion labels, our method predicts temporally dense emotion categories and intensities directly from speech, enabling continuous modulation of facial expressions over time. The diarized emotion signals are encoded as learned embeddings and used to condition a speech-driven 3D animation model based on a hybrid Transformer-Mamba architecture. This design allows effective disentanglement of linguistic content and emotional style while preserving identity and temporal coherence. We evaluate our approach on a large-scale multi-corpus dataset for speech emotion diarization and on the EmoVOCA dataset for emotional 3D facial animation. Quantitative results demonstrate strong frame-level emotion recognition performance and low geometric and temporal reconstruction errors, while qualitative results show smooth emotion transitions and consistent expression control. These findings highlight the effectiveness of frame-level emotion diarization for expressive and controllable 3D talking head generation.

3D人脸动画情感识别语音驱动表情生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。