arXiv:2505.03186cs.SDcs.CV2025-05被引 1

用视听同步学习通用表示,仅用223小时数据实现跨任务强性能。

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

  • 通过对比对齐与生成预测双目标,利用视听自然同步性建模
  • 在LRS2上实现AVSR WER 1.27,噪声下性能提升超70%
  • 适合需要多模态语音处理的科研与工业应用

说话人唇动、语音与语言内容之间的内在同步为提升语音处理任务提供了丰富信息,尤其在传统音频单模态系统失效的挑战场景中。我们提出CoGenAV,一种强大且数据高效的模型,可学习适用于多种语音与视听任务的通用表示。CoGenAV通过优化源自自然视听同步的双重目标(对比特征对齐与生成文本预测)进行训练,仅使用来自LRS2数据集的223小时标注数据。该对比-生成同步策略有效捕捉了跨模态的基本相关性。我们在多个基准测试中展示了所学CoGenAV表示的有效性与通用性:在LRS2上的音视频语音识别(AVSR)任务中达到1.27的词错误率(WER),视觉语音识别(VSR)任务中达20.5的WER;在噪声环境下性能提升超过70%;同时显著改善语音重建任务(如语音增强与分离),并在主动说话人检测(ASD)等视听同步任务中取得竞争性结果。模型将开源,以促进学术界与产业界的进一步发展与协作。

原文摘要 · Abstract (English)

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions where traditional audio-only systems falter. We introduce CoGenAV, a powerful and data-efficient model designed to learn versatile audio-visual representations applicable across a wide range of speech and audio-visual tasks. CoGenAV is trained by optimizing a dual objective derived from natural audio-visual synchrony, contrastive feature alignment and generative text prediction, using only 223 hours of labeled data from the LRS2 dataset. This contrastive-generative synchronization strategy effectively captures fundamental cross-modal correlations. We showcase the effectiveness and versatility of the learned CoGenAV representations on multiple benchmarks. When utilized for Audio-Visual Speech Recognition (AVSR) on LRS2, these representations contribute to achieving a state-of-the-art Word Error Rate (WER) of 1.27. They also enable strong performance in Visual Speech Recognition (VSR) with a WER of 20.5 on LRS2, and significantly improve performance in noisy environments by over 70%. Furthermore, CoGenAV representations benefit speech reconstruction tasks, boosting performance in Speech Enhancement and Separation, and achieve competitive results in audio-visual synchronization tasks like Active Speaker Detection (ASD). Our model will be open-sourced to facilitate further development and collaboration within both academia and industry.

视听融合自监督学习语音识别多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。