LID模型其实靠口音分类,而非语言,新方法提升对口音语音的识别准确率。
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
- 模型误将外语口音当母语,根源是依赖短时语音特征判口音
- 通过分段输入可显著提升对口音语音的识别鲁棒性
- 不依赖单语语音识别系统,适合真实场景中带口音的语音识别
先前研究指出,语言识别(LID)模型在带口音语音上的表现明显下降;但错误的具体原因、程度及特征仍不明确。首先,我们发现一种常见失败模式:LID系统常将第二语言口音语音误判为说话者母语或相关语言。其次,我们提出证据表明,当前最先进模型对短时语音片段的顺序不敏感,说明其主要依赖能表征口音的短时音位特征进行分类,而非语言本身。我们的分析揭示了一种简单有效的改进方法——输入分块处理,可增强模型对口音的鲁棒性。第三,我们提出一种无需依赖单语语音识别(ASR)系统的序列级信息融合策略,有效降低口音与语言混淆,显著提升在带口音语音上的性能,同时保持标准语言识别的相当水平。
原文摘要 · Abstract (English)
Prior research indicates that LID model performance significantly declines on accented speech; however, the specific causes, extent, and characterization of these errors remain under-explored. (i) We identify a common failure mode on accented speech whereby LID systems often misclassify L2 accented speech as the speaker's native language or a related language. (ii) We present evidence suggesting that state-of-the-art models are invariant to permutations of short spans of speech, implying they classify on the basis of short phonotactic features indicative of accent rather than language. Our analysis reveals a simple method to enhance model robustness to accents through input chunking. (iii) We present an approach that integrates sequence-level information into our model without relying on monolingual ASR systems; this reduces accent-language confusion and significantly enhances performance on accented speech while maintaining comparable results on standard LID.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。