arXiv:2409.02657cs.CVcs.AI2024-09被引 5

用文本和音频控制头姿,生成口型同步的自然说话头像视频。

PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation

  • 通过文本与音频联合生成动作潜空间,实现自由头姿控制。
  • 双阶段网络结构提升唇动同步效果,口型误差降低超30%。
  • 适合需要精准表情与动作编辑的数字人、虚拟主播应用。

现有音频驱动说话头生成方法虽能根据音频生成头部姿态,但生成的口型与音频不匹配或不可编辑。本文提出PoseTalk系统,可基于文本提示与音频自由生成口型同步的说话头像视频。核心思路是利用头部姿态连接视觉、语言与音频信号:先设计姿态潜扩散模型(PLD),从文本与音频中联合生成运动潜变量;再针对唇部损失占比不足4%导致优化偏移的问题,提出级联式精修策略,由粗略网络生成基础动作,再由精细网络分层优化唇部细节。实验表明,该方法在姿态多样性与真实感上优于仅用文本或音频的方案,生成视频在自然头姿表现上超越现有最优方法。

原文摘要 · Abstract (English)

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editable. In this study, we propose \textbf{PoseTalk}, a THG system that can freely generate lip-synchronized talking head videos with free head poses conditioned on text prompts and audio. The core insight of our method is using head pose to connect visual, linguistic, and audio signals. First, we propose to generate poses from both audio and text prompts, where the audio offers short-term variations and rhythm correspondence of the head movements and the text prompts describe the long-term semantics of head motions. To achieve this goal, we devise a Pose Latent Diffusion (PLD) model to generate motion latent from text prompts and audio cues in a pose latent space. Second, we observe a loss-imbalance problem: the loss for the lip region contributes less than 4\% of the total reconstruction loss caused by both pose and lip, making optimization lean towards head movements rather than lip shapes. To address this issue, we propose a refinement-based learning strategy to synthesize natural talking videos using two cascaded networks, i.e., CoarseNet, and RefineNet. The CoarseNet estimates coarse motions to produce animated images in novel poses and the RefineNet focuses on learning finer lip motions by progressively estimating lip motions from low-to-high resolutions, yielding improved lip-synchronization performance. Experiments demonstrate our pose prediction strategy achieves better pose diversity and realness compared to text-only or audio-only, and our video generator model outperforms state-of-the-art methods in synthesizing talking videos with natural head motions. Project: https://junleen.github.io/projects/posetalk.

说话头生成姿态控制扩散模型多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。