arXiv:2508.14359cs.CV2025-08

用离散化视觉标记实现情绪可控的逼真说话人脸生成

Taming Transformer for Emotion-Controllable Talking Face Generation

  • 将音频分离并量化视频为视觉标记,解耦情绪与身份特征
  • 引入情感锚点表示,融合情绪信息到视觉标记中
  • 基于自回归变换器生成符合情绪条件的连贯人脸视频,适合影视合成场景

说话人脸生成是一项新颖且具有挑战性的任务,旨在根据给定音频生成生动的说话人脸视频。为实现情绪可控的说话人脸生成,现有方法面临两大挑战:一是如何有效建模与特定情绪相关的多模态关系,二是如何利用该关系生成保持身份特征的情绪化视频。本文提出一种离散化方法解决情绪可控说话人脸生成问题。具体地,采用两种预训练策略将音频分解为独立成分,并将视频量化为视觉标记的组合;随后提出情感锚点(EA)表示,将情绪信息融入视觉标记;最后引入自回归变换器,在给定条件下建模视觉标记的全局分布,并预测用于合成操控视频的索引序列。在包含多种情绪音频的MEAD数据集上进行实验,大量实验证明了该方法在定性和定量上的优越性。

原文摘要 · Abstract (English)

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two challenges: One is how to effectively model the multimodal relationship related to the specific emotion, and the other is how to leverage this relationship to synthesize identity preserving emotional videos. In this paper, we propose a novel method to tackle the emotion-controllable talking face generation task discretely. Specifically, we employ two pre-training strategies to disentangle audio into independent components and quantize videos into combinations of visual tokens. Subsequently, we propose the emotion-anchor (EA) representation that integrates the emotional information into visual tokens. Finally, we introduce an autoregressive transformer to model the global distribution of the visual tokens under the given conditions and further predict the index sequence for synthesizing the manipulated videos. We conduct experiments on the MEAD dataset that controls the emotion of videos conditioned on multiple emotional audios. Extensive experiments demonstrate the superiorities of our method both qualitatively and quantitatively.

说话人脸情绪控制变换器视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。