用自监督模型分析语音中的口音感知机制,发现关键发音特征影响判断。
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
- 通过预训练语音表示提取发音特征,研究口音感知的声学基础。
- 口音强度与美式英语基线的距离显著相关,方向符合预期。
- 揭示了可解释的发音特征在口音识别中的重要作用,适合语音认知研究者。
传统口音感知模型低估了听者依赖的发音特征梯度变化的作用。本文研究当前自监督学习(SSL)语音模型如何编码影响音段口音感知的发音特征级变化。聚焦三种音段:唇齿近音、卷舌闪音和卷舌塞音,这些音在印度本土母语者(如印地语母语者)及其他南亚语言母语者的英语中均一致出现。使用CSLU外语口音英语语料库(Lander, 2007),通过Phonet工具提取发音特征概率,并结合Wav2Vec2-BERT(Barrault et al., 2023)和WavLM(Chen et al., 2022)的预训练表示,以及美式英语母语者对口音的评判结果。探针分析显示,口音强度由部分预训练表示特征最佳预测,其中对比美式英语期望与非母语实际发音的感知显著发音特征被赋予更高权重。基于预训练表示的音段距离与美式及印地语英语基线的多元逻辑回归分析揭示,口音强度的几率与基线距离强相关,方向符合预期。结果表明,自监督语音表示在利用可解释发音特征建模口音感知方面具有重要价值。
原文摘要 · Abstract (English)
Traditional models of accent perception underestimate the role of gradient variations in phonological features which listeners rely upon for their accent judgments. We investigate how pretrained representations from current self-supervised learning (SSL) models of speech encode phonological feature-level variations that influence the perception of segmental accent. We focus on three segments: the labiodental approximant, the rhotic tap, and the retroflex stop, which are uniformly produced in the English of native speakers of Hindi as well as other languages in the Indian sub-continent. We use the CSLU Foreign Accented English corpus (Lander, 2007) to extract, for these segments, phonological feature probabilities using Phonet (Vásquez-Correa et al., 2019) and pretrained representations from Wav2Vec2-BERT (Barrault et al., 2023) and WavLM (Chen et al., 2022) along with accent judgements by native speakers of American English. Probing analyses show that accent strength is best predicted by a subset of the segment's pretrained representation features, in which perceptually salient phonological features that contrast the expected American English and realized non-native English segments are given prominent weighting. A multinomial logistic regression of pretrained representation-based segment distances from American and Indian English baselines on accent ratings reveals strong associations between the odds of accent strength and distances from the baselines, in the expected directions. These results highlight the value of self-supervised speech representations for modeling accent perception using interpretable phonological features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。