arXiv:2512.17562cs.SDcs.AI2025-12被引 3

现代医学语音识别模型用降噪反而更差,降噪会破坏关键声学特征。

When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems

  • 在9种噪声下测试4个主流模型,发现降噪后识别错误率更高
  • 所有40种组合中,原始音频的语义词错率均低于降噪后,最高恶化46.6%
  • 适合医疗语音识别部署者,提醒避免无效甚至有害的降噪预处理

语音增强常被认为能提升嘈杂环境下的自动语音识别(ASR)性能。然而,对于在多样化、含噪数据上训练的大规模现代ASR模型,这一效果并不必然成立。本文系统评估了MetricGAN-plus-voicebank降噪方法在四个先进ASR系统(OpenAI Whisper、NVIDIA Parakeet、Google Gemini Flash 2.0、Parrotlet-a)上的表现,使用500段医疗语音录音,在九种噪声条件下进行测试。采用语义词错率(semWER)作为评价指标,该指标对领域特异性归一化进行了调整。结果显示:在所有噪声条件和模型配置下,降噪处理均导致ASR性能下降。原始噪声音频在全部40种测试配置中(4模型×10条件)均优于降噪后音频,语义词错率绝对增加幅度为1.1%至46.6%。这表明现代ASR模型具备足够的内部噪声鲁棒性,而传统语音增强可能移除了对识别至关重要的声学特征。对于在嘈杂临床环境中部署医疗语音记录系统的从业者而言,本研究提示:使用降噪预处理不仅计算资源浪费,还可能损害转录准确性。

原文摘要 · Abstract (English)

Speech enhancement methods are commonly believed to improve the performance of automatic speech recognition (ASR) in noisy environments. However, the effectiveness of these techniques cannot be taken for granted in the case of modern large-scale ASR models trained on diverse, noisy data. We present a systematic evaluation of MetricGAN-plus-voicebank denoising on four state-of-the-art ASR systems: OpenAI Whisper, NVIDIA Parakeet, Google Gemini Flash 2.0, Parrotlet-a using 500 medical speech recordings under nine noise conditions. ASR performance is measured using semantic WER (semWER), a normalized word error rate (WER) metric accounting for domain-specific normalizations. Our results reveal a counterintuitive finding: speech enhancement preprocessing degrades ASR performance across all noise conditions and models. Original noisy audio achieves lower semWER than enhanced audio in all 40 tested configurations (4 models x 10 conditions), with degradations ranging from 1.1% to 46.6% absolute semWER increase. These findings suggest that modern ASR models possess sufficient internal noise robustness and that traditional speech enhancement may remove acoustic features critical for ASR. For practitioners deploying medical scribe systems in noisy clinical environments, our results indicate that preprocessing audio with noise reduction techniques might not just be computationally wasteful but also be potentially harmful to the transcription accuracy.

语音识别降噪危害医疗AIASR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。