探索模块化设计如何提升语音识别在噪声环境下的鲁棒性
An investigation of modularity for noise robustness in conformer-based ASR
- 用固定路由的Conformer模型测试环境感知模块的效果
- 在CHIME数据集上,区分不同噪声环境效果不佳,仅区分有无噪声最优
- 适合部署在已知或未知噪声环境中的大型语音识别系统参考
当前最先进的自动语音识别(ASR)系统在训练环境外表现会下降。尽管某些新环境可能已知,但可用于适应的数据量很少。本文通过实验研究近期提出的模块化形式是否有助于ASR模型适应新声学环境。采用基于Conformer的模型与固定路由机制,结果表明环境感知确实能提升已知环境下的性能。然而,在本研究使用的CHIME数据集上,分类模块难以区分不同噪声环境,而仅区分噪声与纯净语音的配置表现最佳。该发现对在特定环境中部署大型模型(无论是否事先知晓噪声类型)具有明确指导意义。
原文摘要 · Abstract (English)
Whilst state of the art automatic speech recognition (ASR) can perform well, it still degrades when exposed to acoustic environments that differ from those used when training the model. Unfamiliar environments for a given model may well be known a-priori, but yield comparatively small amounts of adaptation data. In this experimental study, we investigate to what extent recent formalisations of modularity can aid adaptation of ASR to new acoustic environments. Using a conformer based model and fixed routing, we confirm that environment awareness can indeed lead to improved performance in known environments. However, at least on the (CHIME) datasets in the study, it is difficult for a classifier module to distinguish different noisy environments, a simpler distinction between noisy and clean speech being the optimal configuration. The results have clear implications for deploying large models in particular environments with or without a-priori knowledge of the environmental noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。