无需掩码的扩散模型实现跨场景通用唇音同步
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
- 用扩散Transformer直接编辑帧,摆脱掩码依赖
- 在真实与AI生成视频中均显著提升同步精度
- 适合需要保持身份一致性的虚拟人/视频生成场景
唇音同步是将说话人的口型动作与语音音频对齐的关键任务,对生成逼真视频内容至关重要。现有方法多依赖参考帧和掩码填充,难以应对身份一致性、姿态变化、面部遮挡及风格化内容等问题。由于音频条件信号弱于视觉信息,原始视频中的嘴形泄露会降低同步质量。本文提出OmniSync,一种适用于多样化视觉场景的通用唇音同步框架。该方法采用无掩码训练范式,利用扩散Transformer实现无需显式掩码的直接帧编辑,支持无限时长推理,同时保持自然面部动态并维持角色身份。推理阶段引入基于流匹配的渐进噪声初始化,确保姿态与身份一致性,并支持精准口部区域编辑。为解决音频条件信号弱的问题,提出动态时空无分类器引导(DS-CFG)机制,自适应调整时间和空间上的引导强度。我们还构建了首个针对AI生成视频唇音同步的评估基准AIGC-LipSync Benchmark。大量实验表明,OmniSync在视觉质量和唇音同步准确率上均显著优于现有方法,在真实世界与AI生成视频中均表现优异。
原文摘要 · Abstract (English)
Lip synchronization is the task of aligning a speaker's lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content. However, existing methods often rely on reference frames and masked-frame inpainting, which limit their robustness to identity consistency, pose variations, facial occlusions, and stylized content. In addition, since audio signals provide weaker conditioning than visual cues, lip shape leakage from the original video will affect lip sync quality. In this paper, we present OmniSync, a universal lip synchronization framework for diverse visual scenarios. Our approach introduces a mask-free training paradigm using Diffusion Transformer models for direct frame editing without explicit masks, enabling unlimited-duration inference while maintaining natural facial dynamics and preserving character identity. During inference, we propose a flow-matching-based progressive noise initialization to ensure pose and identity consistency, while allowing precise mouth-region editing. To address the weak conditioning signal of audio, we develop a Dynamic Spatiotemporal Classifier-Free Guidance (DS-CFG) mechanism that adaptively adjusts guidance strength over time and space. We also establish the AIGC-LipSync Benchmark, the first evaluation suite for lip synchronization in diverse AI-generated videos. Extensive experiments demonstrate that OmniSync significantly outperforms prior methods in both visual quality and lip sync accuracy, achieving superior results in both real-world and AI-generated videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。