arXiv:2603.04710cs.SDcs.AI2026-03被引 1

音频分离虽提升音质,却让零样本语音识别更差。

When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech

  • 用SAM-Audio增强语音后输入Whisper模型
  • 音质提升但英文和孟加拉语识别错误率均上升
  • 适合研究语音增强对零样本识别影响的学者

近年来,自动语音识别(ASR)与语音增强的进步强化了‘音频越干净识别越准’的共识。本文检验该假设在现代零样本ASR系统中的有效性。我们以OpenAI Whisper为模型,评估其在噪声环境下对孟加拉语和英语语音的零样本转录表现,将SAM-Audio作为预处理步骤。五种Whisper变体在两个数据集上测试:英文数据集上,平均PSNR从32.28 dB升至35.99 dB,71.84%语音段的PSNR提高;但所有配置下字错误率(WER)和字符错误率(CER)均上升。孟加拉语数据集中,Whisper large-v3的WER从65.83%增至77.35%,CER从24.13%升至34.74%;英文数据集上,Whisper base的WER从10.53%升至21.66%,CER从4.48%增至12.50%。逐句分析显示,性能下降影响了大量样本,程度因模型而异。结果表明,信号质量提升未必带来更好识别效果,去噪反而可能降低零样本识别精度。

原文摘要 · Abstract (English)

Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lead to more accurate transcription. In this work, we examine whether this assumption holds for modern zero-shot ASR systems. We conduct a structured empirical study of SAM-Audio as a preprocessing step for zero-shot transcription with OpenAI Whisper. Five Whisper variants are evaluated on noisy Bengali and English speech datasets. On the English dataset, SAM-Audio increases the average PSNR from 32.28 dB to 35.99 dB and achieves higher PSNR for 71.84% of the utterances. However, WER and CER increase in every evaluated model-dataset configuration. On the Bengali dataset, Whisper large-v3 WER increases from 65.83% to 77.35%, while CER increases from 24.13% to 34.74%. On the English dataset, Whisper base WER increases from 10.53% to 21.66%, while CER increases from 4.48% to 12.50%. Utterance-level analysis further shows that the degradation affects a substantial portion of the evaluated samples, although its severity varies across Whisper variants. These findings demonstrate that improved signal-level quality does not necessarily lead to better zero-shot ASR performance and that denoising can reduce recognition accuracy.

语音识别去噪零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。