arXiv:2409.16399cs.SDcs.CL2024-09

用仿生声学特征提升语音识别在噪声和攻击下的鲁棒性

Revisiting Acoustic Features for Robust ASR

  • 基于生物听觉机制设计新声学特征,模拟频率掩蔽与侧抑制
  • DoGSpec在对抗攻击下显著优于主流LogMelSpec,准确率损失小
  • 适合追求高鲁棒性的语音识别系统研发者

自动语音识别(ASR)系统需应对真实环境中多样的噪声,包括环境噪声、混响、特殊音效以及恶意攻击(对抗攻击)。近期工作通过构建新型深度神经网络并扩充训练数据来提升准确率与鲁棒性,但依赖相对简单的声学特征。此类方法虽能增强对训练数据中噪声的适应性,却难以应对未见噪声,且对对抗攻击几乎无效。本文重新审视早期基于生物听觉感知设计的声学特征方法。具体而言,评估了多种仿生声学特征的识别准确率与鲁棒性。除已有特征如伽马时频滤波器组特征(GammSpec)外,还提出两种新特征:频率掩蔽谱图(FreqMask)与伽马时频差谱图(DoGSpec),分别模拟频率掩蔽与侧抑制现象。在多个模型与数据集上的实验表明:(1) DoGSpec相比广泛使用的对数梅尔谱图(LogMelSpec)显著提升鲁棒性,同时准确率下降微小;(2) GammSpec在非对抗性噪声上表现更优,但面对对抗攻击时不如DoGSpec。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems must be robust to the myriad types of noises present in real-world environments including environmental noise, room impulse response, special effects as well as attacks by malicious actors (adversarial attacks). Recent works seek to improve accuracy and robustness by developing novel Deep Neural Networks (DNNs) and curating diverse training datasets for them, while using relatively simple acoustic features. While this approach improves robustness to the types of noise present in the training data, it confers limited robustness against unseen noises and negligible robustness to adversarial attacks. In this paper, we revisit the approach of earlier works that developed acoustic features inspired by biological auditory perception that could be used to perform accurate and robust ASR. In contrast, Specifically, we evaluate the ASR accuracy and robustness of several biologically inspired acoustic features. In addition to several features from prior works, such as gammatone filterbank features (GammSpec), we also propose two new acoustic features called frequency masked spectrogram (FreqMask) and difference of gammatones spectrogram (DoGSpec) to simulate the neuro-psychological phenomena of frequency masking and lateral suppression. Experiments on diverse models and datasets show that (1) DoGSpec achieves significantly better robustness than the highly popular log mel spectrogram (LogMelSpec) with minimal accuracy degradation, and (2) GammSpec achieves better accuracy and robustness to non-adversarial noises from the Speech Robust Bench benchmark, but it is outperformed by DoGSpec against adversarial attacks.

语音识别鲁棒性仿生特征对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。