arXiv:2506.04635cs.CLcs.CV2025-06中稿 · Interspeech 2025

自动构建越语音视频数据集,提升嘈杂环境下的语音识别性能

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

  • 从原始视频自动提取音视频数据,优化采集效率
  • 新数据集使模型在嘈杂环境中表现显著优于纯音频方案
  • 为资源匮乏语言提供可扩展的多模态数据构建路径

音视频语音识别(AVSR)因对噪声具有鲁棒性而受到关注,能有效克服仅依赖音频特征的传统语音识别系统在噪声环境下的局限。然而,多数非英语语言仍受限于缺乏大规模数据集。本文提出一种实用方法,从原始视频中自动化生成AVSR数据,改进现有技术以提升效率与可及性。通过构建越南语基准AVSR模型验证其普适性,实验表明,该自动采集的数据集支持模型在干净环境下表现媲美强健的纯语音识别系统,并在鸡尾酒会等嘈杂场景中显著超越后者。该高效方法为将AVSR推广至更多语言,特别是低资源语言,提供了可行路径。

原文摘要 · Abstract (English)

Audio-Visual Speech Recognition (AVSR) has gained significant attention recently due to its robustness against noise, which often challenges conventional speech recognition systems that rely solely on audio features. Despite this advantage, AVSR models remain limited by the scarcity of extensive datasets, especially for most languages beyond English. Automated data collection offers a promising solution. This work presents a practical approach to generate AVSR datasets from raw video, refining existing techniques for improved efficiency and accessibility. We demonstrate its broad applicability by developing a baseline AVSR model for Vietnamese. Experiments show the automatically collected dataset enables a strong baseline, achieving competitive performance with robust ASR in clean conditions and significantly outperforming them in noisy environments like cocktail parties. This efficient method provides a pathway to expand AVSR to more languages, particularly under-resourced ones.

音视频识别多模态越南语数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。