arXiv:2604.21628cs.SD2026-04中稿 · IEEE ICASSP 2026

分析语音模型中哪部分信息对发音障碍特征最有用

Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0

论文配图:Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
图 1 · 摘自论文原文
  • 用注意力统计池化比较不同层和时间点的语音表示
  • 发音清晰度靠深层特征,辅音不准等依赖时间序列特征
  • 适合语音病理分析、模型可解释性研究者阅读

Wav2vec 2.0(W2V2)在病理语音分析中表现出色,能有效捕捉异常语音特征。然而,其学习表征中哪些组件对特定下游任务最有效仍不明确。本研究基于语音可访问性项目数据集,通过回归分析五类发音障碍描述符:发音清晰度、辅音不准确、不当沉默、声音粗糙和单音调。使用基于W2V2的特征提取器,系统比较了层间与时间维度的聚合策略,采用注意力统计池化。结果表明,发音清晰度最佳由层间表示捕捉,而辅音不准确、声音粗糙和单音调则受益于时间维度建模;对于不当沉默,两种方法无明显优劣。

原文摘要 · Abstract (English)

Wav2vec 2.0 (W2V2) has shown strong performance in pathological speech analysis by effectively capturing the characteristics of atypical speech. Despite its success, it remains unclear which components of its learned representations are most informative for specific downstream tasks. In this study, we address this question by investigating the regression of dysarthric speech descriptors using annotations from the Speech Accessibility Project dataset. We focus on five descriptors, each addressing a different aspect of speech or voice production: intelligibility, imprecise consonants, inappropriate silences, harsh voice and monoloudness. Speech representations are derived from a W2V2-based feature extractor, and we systematically compare layer-wise and time-wise aggregation strategies using attentive statistics pooling. Our results show that intelligibility is best captured through layer-wise representations, whereas imprecise consonants, harsh voice and monoloudness benefit from time-wise modeling. For inappropriate silences, no clear advantage could be observed for either approach.

语音分析可解释性病理语音Wav2vec2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。