arXiv:2502.12414cs.CL2025-02ACL被引 23

发现语音模型在噪声和分布偏移下易产生幻觉,提出新指标HER评估风险。

Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

  • 引入幻觉误差率HER,量化语音模型生成虚假内容的概率。
  • 低词错率可能隐藏高幻觉风险,合成噪声显著提升幻觉率。
  • 适合医疗、法律等高风险领域使用语音模型的开发者参考。

大规模语音基础模型在多任务语音处理中表现优异,但其性能评估仍面临挑战。传统指标如词错误率(WER)和字符错误率(CER)难以反映关键场景下的转录质量,尤其在检测伪造输出方面存在缺陷。这种现象称为幻觉,尤其在医疗、法律、航空等高风险领域后果严重。本文研究了语音识别模型中的幻觉问题,提出幻觉误差率(HER)以量化该现象。对20余种语音模型的分析显示:(1)高WER可能掩盖低幻觉率,而低WER可能隐藏危险幻觉;(2)合成噪声(包括对抗性噪声及白噪声、音高偏移、时间拉伸等常见扰动)会显著增加HER;(3)分布偏移与HER高度相关(α=0.91)。结果表明,在高风险场景中应结合HER与传统指标共同评估模型性能。

原文摘要 · Abstract (English)

Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple speech tasks, including automatic speech recognition (ASR). These models transcend language and domain barriers, yet effectively measuring their performance remains a challenge. Traditional metrics like word error rate (WER) and character error rate (CER) are commonly used to evaluate ASR performance but often fail to reflect transcription quality in critical contexts, particularly when detecting fabricated outputs. This phenomenon, known as hallucination, is especially concerning in high-stakes domains such as healthcare, legal, and aviation, where errors can have severe consequences. In our work, we address this gap by investigating hallucination in ASR models. We examine how factors such as distribution shifts, model size, and model architecture influence the hallucination error rate (HER), a metric we introduce to quantify hallucinations. Our analysis of over 20 ASR models reveals \numinsights~key insights: (1) High WERs can mask low hallucination rates, while low WERs may conceal dangerous hallucinations. (2) Synthetic noise, both adversarial and common perturbations like white noise, pitch shift, and time stretching, increase HER. (3) Distribution shift correlates strongly with HER ($α= 0.91$). Our findings highlight the importance of incorporating HER alongside traditional metrics like WER to better assess ASR model performance, particularly in high-stakes domains.

语音识别幻觉检测模型评估高风险应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。