arXiv:2606.23048cs.SDcs.AI2026-06中稿 · Interspeech 2026被引 3

首个真实语音中幻觉数据集,助力精准检测ASR错误

HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

论文配图:HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems
图 1 · 摘自论文原文
  • 基于真实财报电话录音,人工标注七款主流ASR的幻觉片段
  • 发现幻觉在低误码率语音中仍高频出现,跨模型词汇重合度高
  • 提出新基准测试,现有方法检测效果仅53.1% F1,仍有巨大提升空间

端到端自动语音识别(ASR)系统在自然语音上会产生幻觉,但现有缓解方法通常在非语音或人为损坏音频上评估。我们引入了HALAS,首个由人类标注的真实语音幻觉数据集,涵盖七款先进ASR模型在真实未经处理的财报电话会议录音上的表现。HALAS提供片段级标签,支持对幻觉模式及其严重程度的分析。分析显示,跨模型词汇存在显著重叠,并证实即使在低词错误率(WER)的正确转录中,幻觉仍普遍存在。基于HALAS提出的基准表明,常用字符和语义级指标作为幻觉检测代理时达到81%的ROC-AUC,而最先进检测方法的F1分数仅为53.1%。因此,HALAS建立了首个严谨的非人为伪造的ASR幻觉检测与缓解基准。

原文摘要 · Abstract (English)

End-to-end Automatic Speech Recognition (ASR) systems hallucinate on natural speech, yet existing mitigation methods are typically evaluated on non-speech or artificially corrupted audio. We introduce HALAS, the first human-annotated dataset of naturally occurring hallucinations from seven state-of-the-art ASR models on real unprocessed earnings call recordings. HALAS provides span-level labels, enabling analysis of hallucination patterns and their severity. Our analysis reveals strong cross-model vocabulary overlap and confirms that hallucinations also occur for almost correctly transcribed speech (characterized by a low Word Error Rate). The proposed benchmark with HALAS shows that the character and semantic-level metrics used as a proxy for hallucination detection reach 81% ROC-AUC, while state-of-the-art detection methods achieve an F1 score of only 53.1%. As such, HALAS establishes the first rigorous non-artificial benchmark for the detection and mitigation of ASR hallucinations.

ASR幻觉语音识别数据集评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。