让语音驱动的3D人脸动画实现连续情绪调节,告别离散情绪标签。
EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing
- 通过边界感知嵌入构建连续表情空间,支持平滑情绪调整。
- 在保持口型同步的前提下,情绪表达更自然且可控性更强。
- 适合需要精细情感控制的虚拟人、动画制作和交互应用。
语音驱动的3D人脸动画旨在直接从音频生成逼真且富有表现力的面部运动。尽管近期方法在口型同步上已达到高质量,但普遍依赖离散情绪类别,限制了连续且细粒度的情感控制。本文提出EditEmoTalk,一种支持连续情绪编辑的可控语音驱动3D人脸动画框架。核心思想是引入边界感知语义嵌入,学习情绪间决策边界的正常方向,从而构建连续的表情流形,实现平滑的情绪操作。此外,我们设计了一种情感一致性损失,通过映射网络强制生成运动动态与目标情绪嵌入之间的语义对齐,确保情感表达的忠实性。大量实验表明,EditEmoTalk在可控性、表现力和泛化能力方面均优于现有方法,同时保持准确的口型同步。代码与预训练模型将公开发布。
原文摘要 · Abstract (English)
Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting continuous and fine-grained emotional control. We present EditEmoTalk, a controllable speech-driven 3D facial animation framework with continuous emotion editing. The key idea is a boundary-aware semantic embedding that learns the normal directions of inter-emotion decision boundaries, enabling a continuous expression manifold for smooth emotion manipulation. Moreover, we introduce an emotional consistency loss that enforces semantic alignment between the generated motion dynamics and the target emotion embedding through a mapping network, ensuring faithful emotional expression. Extensive experiments demonstrate that EditEmoTalk achieves superior controllability, expressiveness, and generalization while maintaining accurate lip synchronization. Code and pretrained models will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。