arXiv:2409.03270cs.CV2024-09被引 3

让虚拟人物说话更生动,通过风格建模提升视频多样性。

SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model

  • 用高斯分布建模人脸表情与音频的内在风格
  • 通过对比学习捕捉动态风格,生成更丰富的视频内容
  • 适配稳定扩散模型,实现风格可控的高质量人脸说话视频

说话头生成(THG)通常由音频驱动,是数字人、影视制作和虚拟现实等领域的关键挑战。基于扩散模型的方法虽能生成高质量内容,但常忽略人物固有的个性化风格,如说话习惯和面部表情,导致生成视频缺乏多样性和生动性。为此,本文提出风格增强的生动人像生成框架(SVP),充分挖掘风格信息。首先引入概率风格先验学习,利用人脸表情与音频嵌入建模风格为高斯分布,并通过定制对比目标有效捕捉每段视频的动态风格;随后微调预训练的Stable Diffusion模型,通过交叉注意力将学习到的风格作为控制信号注入生成过程。实验表明,该方法可生成多样化、生动且高质量的视频,且对内在风格具有灵活控制能力,优于现有最先进方法。

原文摘要 · Abstract (English)

Talking Head Generation (THG), typically driven by audio, is an important and challenging task with broad application prospects in various fields such as digital humans, film production, and virtual reality. While diffusion model-based THG methods present high quality and stable content generation, they often overlook the intrinsic style which encompasses personalized features such as speaking habits and facial expressions of a video. As consequence, the generated video content lacks diversity and vividness, thus being limited in real life scenarios. To address these issues, we propose a novel framework named Style-Enhanced Vivid Portrait (SVP) which fully leverages style-related information in THG. Specifically, we first introduce the novel probabilistic style prior learning to model the intrinsic style as a Gaussian distribution using facial expressions and audio embedding. The distribution is learned through the 'bespoked' contrastive objective, effectively capturing the dynamic style information in each video. Then we finetune a pretrained Stable Diffusion (SD) model to inject the learned intrinsic style as a controlling signal via cross attention. Experiments show that our model generates diverse, vivid, and high-quality videos with flexible control over intrinsic styles, outperforming existing state-of-the-art methods.

说话头生成扩散模型风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。