arXiv:2409.05089cs.CV2024-09

用WaveNet和LSTM生成能真实回应说话者的听者头部动作视频

Leveraging WaveNet for Dynamic Listening Head Modeling from Speech

  • 结合WaveNet与LSTM的序列到序列模型生成听者反应
  • 在ViCo数据集上超越基线模型,提升反应真实性
  • 适合影视动画、虚拟会议等需要互动反馈的场景

生成听者面部反应旨在模拟面对面对话中听者的交互反馈。本文目标是通过结合WaveNet与长短期记忆网络(LSTM)的序列到序列模型,生成可信的听者头部视频,使其能真实响应单一说话者。方法注重捕捉听者反馈的细微差别,在保持个体身份特征的同时表达恰当的态度与观点。实验结果表明,该方法在ViCo基准数据集上的表现优于基线模型。

原文摘要 · Abstract (English)

The creation of listener facial responses aims to simulate interactive communication feedback from a listener during a face-to-face conversation. Our goal is to generate believable videos of listeners' heads that respond authentically to a single speaker by a sequence-to-sequence model with an combination of WaveNet and Long short-term memory network. Our approach focuses on capturing the subtle nuances of listener feedback, ensuring the preservation of individual listener identity while expressing appropriate attitudes and viewpoints. Experiment results show that our method surpasses the baseline models on ViCo benchmark Dataset.

人脸生成语音驱动序列建模WaveNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。