arXiv:2502.11515cs.CV2025-02被引 14

用音频直接生成逼真口型动作,支持任意角色说话

SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion

  • 基于条件视频扩散模型,直接从音频生成口型
  • 口型与牙齿同步性更好,能适配未见角色
  • 无需中间表示或额外监督,适合动画角色生成

近年来,扩散模型在音频驱动口型同步方面取得显著进展。然而,现有方法通常依赖受限的音视频对齐先验或多阶段学习中间表示,导致训练流程复杂且运动自然度有限。本文提出 SayAnything,一种条件视频扩散框架,可直接从音频输入合成口型动作并保留说话人身份。我们设计了三个专用模块:身份保持模块、音频引导模块和编辑控制模块。新颖的设计有效平衡潜在空间中的多种条件信号,实现外观、运动及区域特异性生成的精确控制,无需额外监督信号或中间表示。大量实验表明,SayAnything 生成的视频高度真实,口型与牙齿同步性更优,能够使未见过的角色说出任意内容,并有效泛化至动画角色。

原文摘要 · Abstract (English)

Recent advances in diffusion models have led to significant progress in audio-driven lip synchronization. However, existing methods typically rely on constrained audio-visual alignment priors or multi-stage learning of intermediate representations to force lip motion synthesis. This leads to complex training pipelines and limited motion naturalness. In this paper, we present SayAnything, a conditional video diffusion framework that directly synthesizes lip movements from audio input while preserving speaker identity. Specifically, we propose three specialized modules including identity preservation module, audio guidance module, and editing control module. Our novel design effectively balances different condition signals in the latent space, enabling precise control over appearance, motion, and region-specific generation without requiring additional supervision signals or intermediate representations. Extensive experiments demonstrate that SayAnything generates highly realistic videos with improved lip-teeth coherence, enabling unseen characters to say anything, while effectively generalizing to animated characters.

视频生成扩散模型口型同步音频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。