arXiv:2504.07229cs.CLeess.AS2025-04EMNLP被引 1

利用环境视觉信息提升嘈杂场景下的语音识别准确率

Visual-Aware Speech Recognition for Noisy Scenarios

  • 通过多头注意力连接预训练音视频编码器,关联噪声源与视觉线索
  • 在嘈杂场景下相比纯音频模型,识别准确率显著提升
  • 无需说话人可见,适用于更广泛的现实场景,适合语音增强应用

人类在嘈杂环境中能借助口型动作和视觉场景等视觉线索增强听觉感知。然而,当前自动语音识别(ASR)或视听语音识别(AVSR)模型在噪声环境下表现不佳。为此,我们提出一种新模型,通过将噪声源与视觉线索相关联来提升转录效果。不同于依赖说话人唇动且要求可见性的方法,本工作利用环境中的广泛视觉信息,使模型能像人一样自然地从噪声中过滤出语音并改进转录。方法采用预训练的语音与视觉编码器,通过多头注意力机制进行融合,实现视频输入中语音转录与噪声标签预测。我们构建了一个可扩展的数据集生成流程,确保视觉线索与音频噪声相关联。实验表明,在嘈杂场景下,该模型显著优于现有纯音频模型,验证了视觉线索对提升转录准确性的关键作用。

原文摘要 · Abstract (English)

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech Recognition (AVSR) models often struggle in noisy scenarios. To solve this task, we propose a model that improves transcription by correlating noise sources to visual cues. Unlike works that rely on lip motion and require the speaker's visibility, we exploit broader visual information from the environment. This allows our model to naturally filter speech from noise and improve transcription, much like humans do in noisy scenarios. Our method re-purposes pretrained speech and visual encoders, linking them with multi-headed attention. This approach enables the transcription of speech and the prediction of noise labels in video inputs. We introduce a scalable pipeline to develop audio-visual datasets, where visual cues correlate to noise in the audio. We show significant improvements over existing audio-only models in noisy scenarios. Results also highlight that visual cues play a vital role in improved transcription accuracy.

语音识别视听融合噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。