arXiv:2508.08967cs.SDcs.AI2025-08中稿 · IEEE ASRU 2025

提出音频通道归一化方法,提升ASR在不同录音条件下的鲁棒性。

Revealing the Role of Audio Channels in ASR Performance Degradation

  • 通过对比干净参考通道对齐模型内部特征,减少通道差异影响。
  • 在未见通道和语言上显著提升识别准确率,跨域泛化能力强。
  • 适合需要高鲁棒性的语音识别部署场景,如移动端或杂音环境。

预训练自动语音识别(ASR)模型在多种任务中表现优异,但在输入音频来自不同录音通道时性能可能显著下降。以往研究多将此现象归因于训练与测试语料的不匹配,本研究认为,不同通道引起的语音特征变化本身会从根本上损害ASR性能。为此,我们提出一种归一化技术,通过将ASR模型内部特征表示与干净参考通道的特征对齐,减轻通道差异的影响。该方法显著提升了模型在未曾见过的通道和语言上的表现,验证了其在通道与语言差异下的强泛化能力。

原文摘要 · Abstract (English)

Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences.

语音识别通道鲁棒性特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。