arXiv:2409.17674cs.CV2024-09被引 1

用自监督隐空间偏差生成更真实的口说手势视频

Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation

  • 通过扩散模型捕捉隐空间中的像素级运动偏差
  • 在FGD、PSNR等指标上优于当前最优方法,提升最高达8.1%
  • 适合研究语音同步手势生成与自监督表征的学者

手势在协同言语交流中至关重要。现有工作多集中于点级运动变换或全监督运动表示的数据驱动方法,本文聚焦于口说手势的表征,提出一种基于自监督隐空间偏差与像素级运动偏差的方法,并结合扩散模型引入隐运动特征。该方法利用自监督的隐空间偏差来促进手部手势生成,对生成真实手势视频具有关键作用。首次实验结果显示,本方法在生成视频质量上显著提升:FGD、DIV和FVD指标提高2.7至4.5%,PSNR提升8.1%,SSIM提升2.5%,优于当前最先进方法。

原文摘要 · Abstract (English)

Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations through data-driven approaches, we explore the representation of gestures in co-speech, with a focus on self-supervised representation and pixel-level motion deviation, utilizing a diffusion model which incorporates latent motion features. Our approach leverages self-supervised deviation in latent representation to facilitate hand gestures generation, which are crucial for generating realistic gesture videos. Results of our first experiment demonstrate that our method enhances the quality of generated videos, with an improvement from 2.7 to 4.5% for FGD, DIV, and FVD, and 8.1% for PSNR, 2.5% for SSIM over the current state-of-the-art methods.

手势生成扩散模型自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。