arXiv:2409.19501cs.SDcs.AI2024-09被引 3

让说话头像随情绪强度动态变化,更真实自然。

Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation

论文配图:Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
图 1 · 摘自论文原文
  • 用连续潜空间编码情绪类型和强度,实现精细控制。
  • 通过语音语调预测情绪强度,无需逐帧标注。
  • 生成结果更富表现力,适合影视与虚拟人应用。

人类情感表达本质上是动态、复杂且流动的,体现在言语交流中情绪强度的平滑变化。然而,以往基于音频的说话头像生成方法大多忽略了这种强度波动,导致输出情感僵化。本文探索了说话过程中情绪强度的变化规律,提出一种捕捉并生成这些细微波动的方法。具体而言,构建了一个能够生成多种情绪且精确控制强度水平的说话头像框架。该框架通过学习一个连续的情绪潜空间,将情绪类型编码在潜变量方向上,情绪强度则由潜变量的范数体现。此外,为捕捉动态强度变化,采用基于语音到强度的预测器,并利用无情感依赖的强度伪标签方法获取训练信号,无需逐帧标注情绪强度。大量实验与分析验证了所提方法在准确捕捉和再现情绪强度波动方面的有效性,显著提升了生成结果的表现力与真实性。

原文摘要 · Abstract (English)

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by previous audio-driven talking-head generation methods, which often results in static emotional outputs. In this paper, we explore how emotion intensity fluctuates during speech, proposing a method for capturing and generating these subtle shifts for talking-head generation. Specifically, we develop a talking-head framework that is capable of generating a variety of emotions with precise control over intensity levels. This is achieved by learning a continuous emotion latent space, where emotion types are encoded within latent orientations and emotion intensity is reflected in latent norms. In addition, to capture the dynamic intensity fluctuations, we adopt an audio-to-intensity predictor by considering the speaking tone that reflects the intensity. The training signals for this predictor are obtained through our emotion-agnostic intensity pseudo-labeling method without the need of frame-wise intensity labeling. Extensive experiments and analyses validate the effectiveness of our proposed method in accurately capturing and reproducing emotion intensity fluctuations in talking-head generation, thereby significantly enhancing the expressiveness and realism of the generated outputs.

说话头像情绪建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。