arXiv:2510.16567cs.CLcs.SD2025-10被引 7

首个系统评估语音模型幻觉的基准框架,能识别不同类型的错误。

Hallucination Benchmark for Speech Foundation Models

  • 按词汇、语音、形态、语义四维度分类评估语音识别幻觉
  • 在高错误率下仍能区分传统指标忽略的细微错误模式
  • 适合医疗、法律等高风险场景的模型可靠性评估

语音识别系统中的幻觉指神经网络生成的流畅转录与实际语音信号完全无关。这类错误虽看似通顺,却可能误导后续处理,在医疗、法律等关键领域带来严重风险。传统评估指标仅关注错误率,难以区分语音错识与幻觉。为此,我们提出SHALLOW——首个系统性分类与量化语音识别幻觉的基准框架,涵盖词汇、语音、形态、语义四个互补维度,并定义相应指标以生成可解释的模型行为画像。在多种架构与语音领域上评估发现,当识别质量高(低词错误率)时,SHALLOW指标与词错误率(WER)相关性强;但随着WER升高,相关性显著减弱。因此,SHALLOW能捕捉到传统指标在复杂或低质量条件下无法分辨的细粒度错误模式,支持精准诊断模型缺陷并提供改进反馈。

原文摘要 · Abstract (English)

Hallucinations in automatic speech recognition (ASR) systems refer to fluent and coherent transcriptions produced by neural ASR models that are completely unrelated to the underlying acoustic input (i.e., the speech signal). While similar to conventional decoding errors in potentially compromising the usability of transcriptions for downstream applications, hallucinations can be more detrimental due to their preservation of syntactically and semantically plausible structure. This apparent coherence can mislead subsequent processing stages and introduce serious risks, particularly in critical domains such as healthcare and law. Conventional evaluation metrics are primarily centered on error-based metrics and fail to distinguish between phonetic inaccuracies and hallucinations. Consequently, there is a critical need for new evaluation frameworks that can effectively identify and assess models with a heightened propensity for generating hallucinated content. To this end, we introduce SHALLOW, the first benchmark framework that systematically categorizes and quantifies hallucination phenomena in ASR along four complementary axes: lexical, phonetic, morphological, and semantic. We define targeted metrics within each category to produce interpretable profiles of model behavior. Through evaluation across various architectures and speech domains, we have found that SHALLOW metrics correlate strongly with word error rate (WER) when recognition quality is high (i.e., low WER). Still, this correlation weakens substantially as WER increases. SHALLOW, therefore, captures fine-grained error patterns that WER fails to distinguish under degraded and challenging conditions. Our framework supports specific diagnosis of model weaknesses and provides feedback for model improvement beyond what aggregate error rates can offer.

语音识别幻觉检测评估基准可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。