arXiv:2412.17129eess.AScs.SD2024-12中稿 · ICASSP 2025被引 13

对比人耳听觉,揭示语音识别中视觉信息的真实贡献

Uncovering the Visual Contribution in Audio-Visual Speech Recognition

  • 以人类感知为视角评估视觉信息利用率
  • 发现低错误率不等于高信噪比增益
  • 建议未来研究同时报告错误率和有效信噪比增益

音频-视觉语音识别(AVSR)通过融合听觉与视觉语音线索,提升语音识别系统的准确性和鲁棒性。近期进展表明,相比纯音频系统,AVSR在噪声环境下的表现更优。然而,视觉信息的真实贡献程度,以及现有系统是否充分挖掘了视觉线索,仍不明确。本文从人类语音感知角度出发,采用Auto-AVSR、AVEC和AV-RelScore三个系统,首先在0 dB条件下量化视觉贡献的有效信噪比增益,进而分析视觉信息的时间分布特性与词级信息量。结果表明,低词错误率(WER)并不保证高有效信噪比增益。研究显示当前方法未能充分使用视觉信息,建议未来工作在报告WER的同时,补充有效信噪比增益数据。

原文摘要 · Abstract (English)

Audio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR have improved performance in noisy environments compared to audio-only counterparts. However, the true extent of the visual contribution, and whether AVSR systems fully exploit the available cues in the visual domain, remains unclear. This paper assesses AVSR systems from a different perspective, by considering human speech perception. We use three systems: Auto-AVSR, AVEC and AV-RelScore. We first quantify the visual contribution using effective SNR gains at 0 dB and then investigate the use of visual information in terms of its temporal distribution and word-level informativeness. We show that low WER does not guarantee high SNR gains. Our results suggest that current methods do not fully exploit visual information, and we recommend future research to report effective SNR gains alongside WERs.

语音识别多模态视觉线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。