让语音驱动的虚拟人脸表情可控,精确表达情绪。
Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis

- 先学中性表情映射,再用低主成分监督提升情绪控制
- 在保持唇动同步的前提下,情绪分类准确率达新高
- 适合需要精细情感表达的虚拟人、数字主播场景
现有语音驱动人脸生成系统依赖隐式情绪调节,导致控制不精准。训练中若在整个运动空间施加显式情绪损失,会因唇动同步与细粒度情绪控制间的权衡而困难重重。本文发现:虽然情绪线索分布于整个运动空间,但将判别性监督集中在非主成分上可更好平衡情绪与唇动同步,因主成分主要编码高能量的发音和姿态变化。基于此,我们提出Xemo-Talker:首先学习一个中性语音到动作的映射以保证稳定发音与唇动同步,随后引入轻量级情绪分支,通过非主成分子空间监督进行优化。为增强情绪控制,设计三重损失(类间分离、类内紧凑、非主成分对比学习)。输入音频、参考图像和情绪标签后,Xemo-Talker在保持优异唇动同步的同时,实现当前最优的情绪分类准确率,且推理效率高,性能接近真实视频水平。代码已开源。
原文摘要 · Abstract (English)
Precise emotion control in audio-driven talking heads remains a challenge due to the reliance on implicit emotion regulation in existing systems, which often leads to indirect and insufficient control. Additionally, training with explicit emotion-related losses across the entire motion space poses significant difficulties due to the inherent trade-off between accurate lip synchronization and fine-grained emotion control. In this paper, we reveal a key finding: although emotional cues are distributed throughout the motion space, concentrating discriminative supervision on less-principal components achieves a better emotion-lip synchronization balance, as principal components mainly encode high-energy articulation and pose variations. Building on this insight, we propose Xemo-Talker, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision. To enhance emotion control, we design a Tri-Loss consisting of inter-class separation, intra-class compactness, and less-principal contrastive learning. Given an audio input, a reference image, and an emotion label, Xemo-Talker achieves state-of-the-art emotion classification accuracy while maintaining competitive lip synchronization and high inference efficiency, with performance approaching that measured on real videos.The source code is publicly available at https://github.com/chaolongy/Xemo-Talker.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。