新数据集+新方法,让语音识别在嘈杂聚会中更准
Cocktail-Party Audio-Visual Speech Recognition
- 构建含说话与静默面部的1526小时新数据集,模拟真实场景
- 极端噪声下错误率从119%降至39.2%,相对降低67%
- 无需显式分割提示,在复杂环境下仍表现优异
音频-视觉语音识别(AVSR)能有效应对如鸡尾酒会等复杂环境中的语音识别挑战,而仅依赖音频在这些场景中已不可靠。然而,现有AVSR模型多针对理想化场景优化,忽略现实环境中说话与静默面部交替出现的复杂性。本研究提出一个新型音频-视觉鸡尾酒会数据集,用于评估当前AVSR系统并揭示以往方法在真实噪声条件下的局限性。同时,我们构建了一个包含1526小时数据的全新数据集,涵盖说话面部与静默面部片段,显著提升鸡尾酒会环境下的识别性能。所提方法在极端噪声条件下将词错误率(WER)从119%降至39.2%,相对降低67%,且不依赖显式分割信号。
原文摘要 · Abstract (English)
Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AVSR models are often optimized for idealized scenarios with consistently active speakers, overlooking the complexities of real-world settings that include both speaking and silent facial segments. This study addresses this gap by introducing a novel audio-visual cocktail-party dataset designed to benchmark current AVSR systems and highlight the limitations of prior approaches in realistic noisy conditions. Additionally, we contribute a 1526-hour AVSR dataset comprising both talking-face and silent-face segments, enabling significant performance gains in cocktail-party environments. Our approach reduces WER by 67% relative to the state-of-the-art, reducing WER from 119% to 39.2% in extreme noise, without relying on explicit segmentation cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。