arXiv:2509.21887cs.CVcs.MM2025-09

让视频配音更自然真实,能模仿特定说话习惯且不怕遮挡物干扰。

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

  • 基于Stable Diffusion改进,加入唇部习惯建模与遮挡感知训练。
  • 在多个数据集上实现更高嘴型相似度和遮挡鲁棒性,无需昂贵先验。
  • 适合需要高质量语音同步视频生成的研究与工业应用。

视觉配音任务旨在生成与驱动音频同步的口型动作,近年来进展显著。然而两大缺陷制约其广泛应用:(1) 仅依赖音频的驱动方式难以捕捉说话人特有的唇部习惯,导致生成口型与目标角色不符;(2) 传统盲修补方法在处理麦克风、手部等遮挡时易产生视觉伪影,限制实际部署。本文提出StableDub,一种新颖简洁的框架,融合唇习惯感知建模与遮挡鲁棒合成。基于Stable Diffusion主干网络,设计唇习惯调制机制,联合建模音素级音视频同步与说话人特有面部动态。为在遮挡下生成合理唇形与物体外观,引入显式暴露遮挡物的遮挡感知训练策略。结合该设计,模型无需以往方法中的高成本先验,展现出在计算密集的扩散模型主干上的卓越训练效率。为进一步优化架构层面的训练效率,引入混合Mamba-Transformer结构,在低资源场景中表现更优。大量实验表明,StableDub在唇习惯相似度与遮挡鲁棒性方面均优于现有方法,同时在音唇同步、视频质量与分辨率一致性上也领先。本工作从多维度拓展了视觉配音的应用范围,演示视频见https://stabledub.github.io。

原文摘要 · Abstract (English)

The visual dubbing task aims to generate mouth movements synchronized with the driving audio, which has seen significant progress in recent years. However, two critical deficiencies hinder their wide application: (1) Audio-only driving paradigms inadequately capture speaker-specific lip habits, which fail to generate lip movements similar to the target avatar; (2) Conventional blind-inpainting approaches frequently produce visual artifacts when handling obstructions (e.g., microphones, hands), limiting practical deployment. In this paper, we propose StableDub, a novel and concise framework integrating lip-habit-aware modeling with occlusion-robust synthesis. Specifically, building upon the Stable-Diffusion backbone, we develop a lip-habit-modulated mechanism that jointly models phonemic audio-visual synchronization and speaker-specific orofacial dynamics. To achieve plausible lip geometries and object appearances under occlusion, we introduce the occlusion-aware training strategy by explicitly exposing the occlusion objects to the inpainting process. By incorporating the proposed designs, the model eliminates the necessity for cost-intensive priors in previous methods, thereby exhibiting superior training efficiency on the computationally intensive diffusion-based backbone. To further optimize training efficiency from the perspective of model architecture, we introduce a hybrid Mamba-Transformer architecture, which demonstrates the enhanced applicability in low-resource research scenarios. Extensive experimental results demonstrate that StableDub achieves superior performance in lip habit resemblance and occlusion robustness. Our method also surpasses other methods in audio-lip sync, video quality, and resolution consistency. We expand the applicability of visual dubbing methods from comprehensive aspects, and demo videos can be found at https://stabledub.github.io.

视频生成语音同步扩散模型唇语建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。