研究非语音音频如何导致Whisper模型生成错误转录,并提出后处理方案消除幻觉。
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
- 用各类非语音音效诱发Whisper模型幻觉,发现存在高频重复的错误输出模式。
- 通过构建幻觉词袋(BoH)对转录文本后处理,显著降低词错误率(WER)。
- 适合关注ASR鲁棒性、模型幻觉治理的研究者与工程落地开发者。
深度神经网络在自动语音识别(ASR)中的幻觉问题仍是关键挑战。本文研究了在推理过程中引入非语音音频片段时,Whisper ASR模型产生的幻觉现象。通过使用多种声音诱导幻觉,我们发现存在一组频繁出现的特定幻觉内容。随后,我们分析了在语音中加入此类声音所引发的幻觉。最后,我们提出一种‘幻觉词袋’(Bag of Hallucinations, BoH)方法,可通过后处理转录文本有效消除幻觉影响。实验结果表明,该后处理方法能显著降低词错误率(WER),并成为应对有害幻觉的有效防护手段。
原文摘要 · Abstract (English)
Hallucinations of deep neural models are amongst key challenges in automatic speech recognition (ASR). In this paper, we investigate hallucinations of the Whisper ASR model induced by non-speech audio segments present during inference. By inducting hallucinations with various types of sounds, we show that there exists a set of hallucinations that appear frequently. We then study hallucinations caused by the augmentation of speech with such sounds. Finally, we describe the creation of a bag of hallucinations (BoH) that allows to remove the effect of hallucinations through the post-processing of text transcriptions. The results of our experiments show that such post-processing is capable of reducing word error rate (WER) and acts as a good safeguard against problematic hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。