arXiv:2503.16357cs.CVcs.SD2025-03中稿 · ICME 2025被引 6

统一框架提升语音视频同步精度,适配多种视觉表示

UniSync: A Unified Framework for Audio-Visual Synchronization

  • 基于嵌入相似性评估音视频同步,兼容多种音频与视觉特征
  • 引入边界损失和跨说话人不同步对,显著提升判别能力
  • 适用于自然与AI生成视频,尤其适合人脸生成场景

语音视频中的精准音视频同步对内容质量和观众理解至关重要。现有方法虽在规则驱动和端到端学习方面取得进展,但仍依赖有限的音视频表征和次优的学习策略,限制了其在复杂场景下的表现。为此,我们提出UniSync,一种基于嵌入相似性的音视频同步评估新方法。该方法兼容多种音频表示(如梅尔谱图、HuBERT)和视觉表示(如RGB图像、人脸分割图、面部关键点、3DMM),有效处理其显著的维度差异。通过引入基于边界的损失函数和跨说话人不同步样本,增强对比学习框架的判别能力。UniSync在标准数据集上优于现有方法,并展现出在多样化音视频表征下的通用性。将其集成至说话人脸生成框架后,显著提升了自然与AI生成内容的同步质量。

原文摘要 · Abstract (English)

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning techniques. However, these methods often rely on limited audio-visual representations and suboptimal learning strategies, potentially constraining their effectiveness in more complex scenarios. To address these limitations, we present UniSync, a novel approach for evaluating audio-visual synchronization using embedding similarities. UniSync offers broad compatibility with various audio representations (e.g., Mel spectrograms, HuBERT) and visual representations (e.g., RGB images, face parsing maps, facial landmarks, 3DMM), effectively handling their significant dimensional differences. We enhance the contrastive learning framework with a margin-based loss component and cross-speaker unsynchronized pairs, improving discriminative capabilities. UniSync outperforms existing methods on standard datasets and demonstrates versatility across diverse audio-visual representations. Its integration into talking face generation frameworks enhances synchronization quality in both natural and AI-generated content.

音视频同步对比学习人脸生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。