用混合高斯模型捕捉发音变体,提升异常发音识别准确率
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
- 用高斯混合模型建模音素的多个子簇,捕捉发音环境差异
- 在五组数据中四组达到顶尖性能,尤其擅长失语和非母语语音
- 结合自监督模型特征,更有效捕捉发音变异,适合语音病理评估
音位变体指同一音素在不同语音环境下的发音差异。建模音位变体对异常发音评估至关重要,即区分异常与正常发音。然而,现有基于音素分类器的方法常将多种发音形式简化为单一音素,忽略了音位变体的复杂性。受冻结自监督语音模型(S3M)特征的声学建模能力启发,我们提出 MixGoP,利用高斯混合模型对具有多个子簇的音素分布进行建模。实验表明,MixGoP 在五组数据中的四组上达到最先进水平,涵盖失语症语音与非母语语音。分析进一步显示,S3M 特征比 MFCCs 与梅尔谱图更有效地捕捉音位变体,凸显 MixGoP 与 S3M 特征结合的优势。
原文摘要 · Abstract (English)
Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。