arXiv:2412.01064cs.CVcs.AI2024-12ICCV被引 43

用流匹配模型实现语音驱动口型动画,生成更快更连贯。

FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait

  • 采用正交运动潜空间与流匹配机制,提升生成效率与连贯性。
  • 在多个指标上优于当前最佳方法,生成速度更快且动作更自然。
  • 支持情感增强,适合需要表情生动的语音驱动视频场景。

随着基于扩散模型的生成技术快速发展,人物图像动画已取得显著进展。然而,由于其迭代采样特性,仍面临时间一致性视频生成和快速采样难题。本文提出FLOAT,一种基于流匹配生成模型的语音驱动口型动画方法。不同于像素级潜空间,我们采用学习得到的正交运动潜空间,实现高效且可编辑的时间一致运动生成。为此,引入基于Transformer的向量场预测器,并设计有效的帧级条件机制。此外,该方法支持语音驱动的情感增强,使表达性动作自然融入。大量实验表明,本方法在视觉质量、运动保真度和效率方面均超越当前最优的语音驱动口型动画方法。

原文摘要 · Abstract (English)

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative sampling nature. This paper presents FLOAT, an audio-driven talking portrait video generation method based on flow matching generative model. Instead of a pixel-based latent space, we take advantage of a learned orthogonal motion latent space, enabling efficient generation and editing of temporally consistent motion. To achieve this, we introduce a transformer-based vector field predictor with an effective frame-wise conditioning mechanism. Additionally, our method supports speech-driven emotion enhancement, enabling a natural incorporation of expressive motions. Extensive experiments demonstrate that our method outperforms state-of-the-art audio-driven talking portrait methods in terms of visual quality, motion fidelity, and efficiency.

语音驱动口型动画流匹配运动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。