研究语音识别模型在噪声和分布偏移下的鲁棒性,发现针对性训练可提升性能。
Evaluating and Improving the Robustness of Speech Command Recognition Models to Noise and Distribution Shifts
- 通过噪声感知训练提升模型对异常输入的适应能力
- 引入公平性与鲁棒性指标,量化模型在不同场景的表现差异
- 为语音识别模型的泛化能力优化提供实证指导
尽管计算机视觉领域已发现分布内(ID)与分布外(OOD)准确率之间存在强关联,但音频模型中此类关系仍缺乏深入探讨。本文研究了训练条件与输入特征如何影响语音关键词分类器在分布外情况下的鲁棒性与泛化能力。我们在多个评估集上对比了几种神经网络架构。为量化噪声对泛化的影响,采用两个指标:公平性(F),衡量相较于基线模型的整体准确率提升;鲁棒性(R),评估分布内与分布外性能的收敛程度。结果表明,噪声感知训练在某些配置下能有效提升鲁棒性。这些发现揭示了基于噪声增强在语音模型泛化中的优势与局限。
原文摘要 · Abstract (English)
Although prior work in computer vision has shown strong correlations between in-distribution (ID) and out-of-distribution (OOD) accuracies, such relationships remain underexplored in audio-based models. In this study, we investigate how training conditions and input features affect the robustness and generalization abilities of spoken keyword classifiers under OOD conditions. We benchmark several neural architectures across a variety of evaluation sets. To quantify the impact of noise on generalization, we make use of two metrics: Fairness (F), which measures overall accuracy gains compared to a baseline model, and Robustness (R), which assesses the convergence between ID and OOD performance. Our results suggest that noise-aware training improves robustness in some configurations. These findings shed new light on the benefits and limitations of noise-based augmentation for generalization in speech models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。