arXiv:2608.10836cs.SDcs.AI2026-08

让语音大模型学会识别低质量耳语,减少误译。

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

论文配图:Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
图 1 · 摘自论文原文
  • 通过自监督学习让模型感知语音信号缺陷,生成不确定性判断。
  • 在AISHELL6-Whisper数据集上字错误率降低17%,幻觉率从25%降至4.5%。
  • 适合需要高可靠性的耳语识别场景,如医疗、隐私保护应用。

耳语语音信号模糊性导致语音识别系统陷入两种对立的失败模式:无法识别耳语或将噪声误译为文本。本文提出Whisper-Aware LLM框架,使音频大模型具备对不确定性的感知与响应能力。模型通过针对性自监督任务,学习量化声学信号的物理缺陷,形成内在自知能力。该不确定性通过创新的置信度融合解码机制实现应用,为大模型解码器提供高层指令和帧级注意力调制。实验验证了该方法的有效性:在AISHELL6-Whisper数据集上达到新的最先进水平,相对字错误率(CER)降低17%;同时直接缓解可靠性权衡问题,幻觉率从超过25%降至4.5%。

原文摘要 · Abstract (English)

The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio-LLM to perceive and react to this uncertainty. Our model develops an intrinsic self-awareness by learning to quantify the physical deficiencies of acoustic signals through targeted self-supervised tasks. This learned uncertainty is then operationalized via a novel Confidence-Fused Decoding mechanism, which provides both high-level instructions and frame-level attention modulation to the LLM decoder. Our experiments confirm the effectiveness of this approach. The model sets a new state-of-the-art on whispered speech with a 17% relative CER reduction on AISHELL6-Whisper. At the same time, it directly addresses the reliability trade-off, with hallucination rates dropping from over 25% to 4.5%.

语音识别耳语识别自监督学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。