arXiv:2606.23948cs.CL2026-06中稿 · presentation at In…

探究语音模型如何编码非裔美国英语的辅音连缀省略现象

Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English

论文配图:Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English
图 1 · 摘自论文原文
  • 分层探测wav2vec2与Whisper模型内部表示
  • 还原后的辅音仍保留原始发音线索
  • 揭示语音模型对语音变异的结构化编码能力

自监督和监督式语音模型被越来越多地用于探究其内部表征所编码的语言信息及其抽象层次。一个尚未充分研究的现象是非裔美国英语(AAE)中的辅音连缀省略(CCR),这是一种普遍的音系过程,也是自动语音识别(ASR)差异的来源。为分析CCR的表征方式,我们对wav2vec2-base和Whisper-small模型进行了独立说话人分层探测,采用两个任务:音段省略检测和底层连缀身份恢复。两个模型均能以高准确率区分省略形式与标准形式。关键发现是,省略后的音段仍保留其底层塞音的线索,表明CCR在模型中被编码为有结构的音系变异,而非简单的音段删除。这些结果表明,现代语音模型能够结构化地编码AAE CCR模式。

原文摘要 · Abstract (English)

Self-supervised and supervised speech models are increasingly used to investigate which linguistic information their internal representations encode, and at what level of abstraction they encode it. One underexplored phenomenon is consonant cluster reduction (CCR) in African American English (AAE), a widespread phonological process and a source of automatic speech recognition (ASR) disparity. To examine how CCR is represented, we conduct speaker-independent layer-wise probing of wav2vec2-base and Whisper-small using two tasks: segmental reduction detection and segmental restoration of underlying cluster identity. Both models distinguish reduced and canonical forms with high accuracy. Crucially, reduced segments retain cues to their underlying stops, indicating that CCR is encoded as structured gradient phonological variation rather than simple segmental deletion. These results demonstrate structured phonological encoding of AAE CCR patterns in modern speech models.

语音识别音系学语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。