无需遮罩训练,通过两阶段优化实现自然唇形同步与画面稳定。
SyncAnyone: Implicit Disentanglement via Progressive Self-Correction for Lip-Syncing in the wild
- 先用扩散模型修复遮罩后的嘴部动作,精准对齐音频
- 再用合成数据无遮罩微调,消除面部和背景伪影
- 适合真实场景下高质量视频配音,尤其注重画面一致性
高质量AI视频配音需精确的音唇同步、高保真视觉生成以及身份与背景的忠实保留。现有方法多采用遮罩训练策略,在说话人头部视频中掩码口部区域,让模型从受损输入和目标音频中重建唇动。尽管提升音唇同步精度,但破坏时空上下文,导致动态面部动作表现差,且面部结构与背景一致性下降。为此,我们提出SyncAnyone,一种新型两阶段学习框架,同时实现精准运动建模与高视觉保真度。第一阶段训练基于扩散模型的视频变换器,完成遮罩口部修复,利用其强时空建模能力生成音控唇动。由于输入损坏,周边面部及背景可能出现微小伪影。第二阶段设计无遮罩调优流程:基于第一阶段模型,构建数据生成管道,通过源视频与随机采样音频合成伪配对视频样本,并在该合成数据上微调模型,实现精准唇部编辑与更佳背景一致性。大量实验表明,本方法在真实场景下的音唇同步任务中,视觉质量、时间连贯性与身份保留均达当前最优水平。
原文摘要 · Abstract (English)
High-quality AI-powered video dubbing demands precise audio-lip synchronization, high-fidelity visual generation, and faithful preservation of identity and background. Most existing methods rely on a mask-based training strategy, where the mouth region is masked in talking-head videos, and the model learns to synthesize lip movements from corrupted inputs and target audios. While this facilitates lip-sync accuracy, it disrupts spatiotemporal context, impairing performance on dynamic facial motions and causing instability in facial structure and background consistency. To overcome this limitation, we propose SyncAnyone, a novel two-stage learning framework that achieves accurate motion modeling and high visual fidelity simultaneously. In Stage 1, we train a diffusion-based video transformer for masked mouth inpainting, leveraging its strong spatiotemporal modeling to generate accurate, audio-driven lip movements. However, due to input corruption, minor artifacts may arise in the surrounding facial regions and the background. In Stage 2, we develop a mask-free tuning pipeline to address mask-induced artifacts. Specifically, on the basis of the Stage 1 model, we develop a data generation pipeline that creates pseudo-paired training samples by synthesizing lip-synced videos from the source video and random sampled audio. We further tune the stage 2 model on this synthetic data, achieving precise lip editing and better background consistency. Extensive experiments show that our method achieves state-of-the-art results in visual quality, temporal coherence, and identity preservation under in-the wild lip-syncing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。