让人脸表情换情绪但不改变口型,通过对比学习分离内容与情感特征。
Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
- 用对比学习分离语音内容和表情情感特征
- 在多个数据集上实现更精准的表情操控与口型同步
- 适合需要精确控制表情又保留原声口型的研究者
语音保持的人脸表情操控(SPFEM)旨在修改说话人面部以呈现指定参考情绪,同时保留源语音的嘴部动作。参考与源输入中蕴含的情绪和内容信息可提供直接准确的监督信号,但二者在说话过程中存在内在耦合,影响监督效果。本文提出对比解耦表示学习(CDRL)算法,通过对比内容表示学习(CCRL)模块学习音频特征作为内容先验,指导源输入内容表示学习;通过对比情感表示学习(CERL)模块利用预训练视觉-语言模型获取情感先验,引导参考输入情感表示学习。进一步引入情感感知与情感增强的对比学习策略,确保学习到情感无关的内容表示和内容无关的情感表示。训练期间,解耦的表示用于监督生成过程,提升表情操控精度与音视频唇动同步性。大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Speech-preserving facial expression manipulation (SPFEM) aims to modify a talking head to display a specific reference emotion while preserving the mouth animation of source spoken contents. Thus, emotion and content information existing in reference and source inputs can provide direct and accurate supervision signals for SPFEM models. However, the intrinsic intertwining of these elements during the talking process poses challenges to their effectiveness as supervisory signals. In this work, we propose to learn content and emotion priors as guidance augmented with contrastive learning to learn decoupled content and emotion representation via an innovative Contrastive Decoupled Representation Learning (CDRL) algorithm. Specifically, a Contrastive Content Representation Learning (CCRL) module is designed to learn audio feature, which primarily contains content information, as content priors to guide learning content representation from the source input. Meanwhile, a Contrastive Emotion Representation Learning (CERL) module is proposed to make use of a pre-trained visual-language model to learn emotion prior, which is then used to guide learning emotion representation from the reference input. We further introduce emotion-aware and emotion-augmented contrastive learning to train CCRL and CERL modules, respectively, ensuring learning emotion-independent content representation and content-independent emotion representation. During SPFEM model training, the decoupled content and emotion representations are used to supervise the generation process, ensuring more accurate emotion manipulation together with audio-lip synchronization. Extensive experiments and evaluations on various benchmarks show the effectiveness of the proposed algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。