分析癌症患者语音模型如何捕捉声学特征,揭示其关键依赖。
What Does a Pathological Speech Assessment Model Know about Acoustic Features? A Case Study on Oral and Oropharyngeal Cancer Patients

- 用典型相关分析对比模型嵌入与声学特征
- 频谱和语调特征相关性最高,首阶梅尔倒谱系数最强
- 为病理性语音评估提供特征选择参考
本研究通过典型相关分析,探究基于Wav2Vec 2.0的口腔及口咽癌患者语音可懂度评估模型的可解释性。通过测量模型嵌入与eGeMAPS低层描述符(LLDs)之间的相关性,分析声学信息在各网络层的编码方式。分析分两个层面进行:单个LLD逐层分析,以及语调、频谱和声音质量三组特征的群体级分析。结果表明,模型表征与频谱和语调特征相关性最强,第一阶梅尔倒谱系数(MFCC)在所有层中相关性最高。在群体层面,频谱组与语调组的相关性分别为0.77和0.71,声音质量组为0.65。该研究不仅提升模型可解释性,也为病理性语音评估中的声学特征选择提供实用指导。
原文摘要 · Abstract (English)
This work investigates the interpretability of a Wav2Vec 2.0based speech intelligibility assessment model for oral and oropharyngeal cancer patients through canonical correlation analysis. By measuring the correlation between the model embeddings and eGeMAPS low-level descriptors (LLDs) as an interpretable reference, we analyze how acoustic information is encoded across the model layers. The analysis is conducted at two levels: individual LLDs layer-wise, and group-level: prosodic, spectral, and voice quality. Results show that the learned representations are most strongly correlated with spectral and prosodic features, with the first MFCC coefficient yielding the highest correlations across all layers. At the group level, spectral and prosodic groups achieve correlations of 0.77 and 0.71 respectively, while voice quality reaches 0.65. Beyond model interpretability, this work also offers practical guidance on acoustic feature selection for pathological speech assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。