UniSync统一框架提升复杂场景下唇形同步精度与泛化能力
UniSync: Towards Generalizable and High-Fidelity Lip Synchronization for Challenging Scenarios
- 无掩码姿态锚定训练,避免颜色失真
- 微调小样本多样视频后显著提升跨场景适应力
- 支持真人与风格化头像,适合真实应用部署
唇形同步旨在生成与给定音频匹配的逼真说话视频,对高质量视频配音至关重要。现有方法存在根本缺陷:基于掩码的方法易产生局部颜色偏差,而无掩码方法常出现全局背景纹理错位。此外,多数方法在风格化头像、面部遮挡和极端光照等真实场景下表现不佳。本文提出UniSync,一种统一框架,可在多样场景中实现高保真唇形同步。具体而言,UniSync采用无掩码的姿态锚定训练策略,保持头部运动一致性并消除合成颜色伪影;同时结合基于掩码的融合推理,确保结构精确与平滑过渡。值得注意的是,仅在少量多样视频上微调即可赋予模型卓越的领域适应性,有效处理复杂边界情况。我们还构建了RealWorld-LipSync基准测试,涵盖真人与风格化头像等多种应用场景。大量实验表明,UniSync显著优于当前最优方法,推动该领域向真正通用且可投入生产的唇形同步迈进。
原文摘要 · Abstract (English)
Lip synchronization aims to generate realistic talking videos that match given audio, which is essential for high-quality video dubbing. However, current methods have fundamental drawbacks: mask-based approaches suffer from local color discrepancies, while mask-free methods struggle with global background texture misalignment. Furthermore, most methods struggle with diverse real-world scenarios such as stylized avatars, face occlusion, and extreme lighting conditions. In this paper, we propose UniSync, a unified framework designed for achieving high-fidelity lip synchronization in diverse scenarios. Specifically, UniSync uses a mask-free pose-anchored training strategy to keep head motion and eliminate synthesis color artifacts, while employing mask-based blending consistent inference to ensure structural precision and smooth blending. Notably, fine-tuning on compact but diverse videos empowers our model with exceptional domain adaptability, handling complex corner cases effectively. We also introduce the RealWorld-LipSync benchmark to evaluate models under real-world demands, which covers diverse application scenarios including both human faces and stylized avatars. Extensive experiments demonstrate that UniSync significantly outperforms state-of-the-art methods, advancing the field towards truly generalizable and production-ready lip synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。