arXiv:2603.08249eess.AScs.CL2026-03

用合成视频让无视觉数据语言实现高质量语音识别。

Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios with Synthetic Visual Data

  • 用真实音频驱动静态人脸生成合成视频流
  • 在加泰罗尼亚语上达到接近顶尖性能,参数更少
  • 适合资源匮乏语言的多模态语音识别研究

音频视觉语音识别(AVSR)通过结合听觉与视觉线索提升复杂环境下的识别鲁棒性,但多数低资源语言因缺乏标注视频语料而难以应用。本文提出一种零音视频资源的AVSR框架,利用真实音频驱动静态人脸生成合成视觉流。首先在西班牙语基准上评估合成视觉增强效果,随后应用于无标注音视频语料的加泰罗尼亚语。我们合成超过700小时的说话人头部视频,并微调预训练的AV-HuBERT模型。在人工标注的加泰罗尼亚语基准上,该模型实现近状态水平表现,参数量和训练数据远少于现有方法,优于同规模纯音频基线模型,并在噪声下保持多模态优势。可扩展的合成视频为零音视频资源场景下的AVSR提供了可行替代方案。

原文摘要 · Abstract (English)

Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. We propose a zero-AV-resource AVSR framework that relies on synthetic visual streams generated by lip-syncing static facial images with real audio. We first evaluate synthetic visual augmentation on Spanish benchmarks, then apply it to Catalan, a language with no annotated audiovisual corpora. We synthesize over 700 hours of talking-head video and fine-tune a pre-trained AV-HuBERT model. On a manually annotated Catalan benchmark, our model achieves near state-of-the-art performance with much fewer parameters and training data, outperforms an identically trained audio-only baseline, and preserves multimodal advantages in noise. Scalable synthetic video thus offers a viable substitute for real recordings in zero-AV-resource AVSR.

语音识别多模态合成数据低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。