arXiv:2505.00497cs.CV2025-05被引 9

解决高分辨率视频配音中口型同步的漏表达与遮挡问题

KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution

  • 分两阶段设计,用掩码策略防止表情泄露
  • 在唇部重建和跨同步任务上达最新水平
  • 适合自动化配音、虚拟人等真实场景应用

唇同步任务是将现有视频中的嘴型动作与新输入音频对齐,通常被视为语音驱动面部动画的一个简化版本。然而,除了常见的说话头生成问题(如时间一致性),唇同步还面临表情泄露和面部遮挡等新挑战,严重影响自动配音等实际应用,但现有工作常忽略这些问题。为此,我们提出KeySync,一种两阶段框架,在保证时间一致性的同时,通过精心设计的掩码策略解决泄露和遮挡问题。实验表明,KeySync在唇部重建和跨同步任务上达到当前最优效果,视觉质量提升,表情泄露显著减少,依据我们提出的新型泄漏度量指标LipLeak。此外,我们验证了掩码方法在处理遮挡上的有效性,并通过多个消融实验证明了架构设计的合理性。代码与模型权重可在https://antonibigata.github.io/KeySync获取。

原文摘要 · Abstract (English)

Lip synchronization, known as the task of aligning lip movements in an existing video with new input audio, is typically framed as a simpler variant of audio-driven facial animation. However, as well as suffering from the usual issues in talking head generation (e.g., temporal consistency), lip synchronization presents significant new challenges such as expression leakage from the input video and facial occlusions, which can severely impact real-world applications like automated dubbing, but are often neglected in existing works. To address these shortcomings, we present KeySync, a two-stage framework that succeeds in solving the issue of temporal consistency, while also incorporating solutions for leakage and occlusions using a carefully designed masking strategy. We show that KeySync achieves state-of-the-art results in lip reconstruction and cross-synchronization, improving visual quality and reducing expression leakage according to LipLeak, our novel leakage metric. Furthermore, we demonstrate the effectiveness of our new masking approach in handling occlusions and validate our architectural choices through several ablation studies. Code and model weights can be found at https://antonibigata.github.io/KeySync.

唇同步语音驱动表情泄露虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。