不用遮罩也能精准配音,靠生成数据自动生成高质量假配对视频。
From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
- 用掩码修复模型生成伪配对数据,再训练无需遮罩的编辑模型。
- 在多个场景下实现最佳唇形同步与视觉质量,抗遮挡能力更强。
- 适合需要高鲁棒性视频配音的工业应用或影视后期制作。
音驱动视频配音旨在使视频中人物口型与新语音同步,但缺乏理想训练数据——仅口型不同的成对视频。现有方法依赖遮罩修补,但会破坏时空上下文,导致身份漂移和鲁棒性差(如遮挡时),还引发口型泄漏,影响唇形同步。为此,我们提出X-Dub,一种基于扩散变换器的两阶段生成自举框架,实现无遮罩配音。核心思想是将掩码修补模型仅用作专用数据生成器,合成可扩展、高保真的伪配对数据,用于训练并自举一个鲁棒的无遮罩编辑模型作为最终配音器。该模型摆脱遮罩伪影,利用完整视频输入实现高保真推理。我们进一步提出时间步自适应多阶段学习,解耦扩散过程中结构、口型运动和纹理的冲突目标,促进稳定收敛与先进编辑质量。此外,我们构建了X-DubBench基准,涵盖多样场景。大量实验表明,本方法在唇形同步、视觉质量和鲁棒性上均达当前最优水平。
原文摘要 · Abstract (English)
Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spatiotemporal context, leading to identity drift and poor robustness (e.g., to occlusions), while also inducing lip-shape leakage that degrades lip sync. To bridge this gap, we propose X-Dub, a novel two-stage generative bootstrapping framework leveraging powerful Diffusion Transformers to unlock mask-free dubbing. Our core insight is to repurpose a mask-based inpainting model exclusively as a dedicated data generator to synthesize scalable, high-fidelity pseudo-paired data, which is subsequently utilized to train and bootstrap a robust, mask-free editing model as the final video dubber. The final dubber is liberated from masking artifacts and leverages the complete video input for high-fidelity inference. We further introduce timestep-adaptive multi-phase learning to disentangle conflicting objectives (structure, lip motion, and texture) across diffusion phases, facilitating stable convergence and advanced editing quality. Additionally, we present X-DubBench, a benchmark for diverse scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance with superior lip sync, visual quality, and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。