用Wav2Vec2的语音识别能力评估头颈癌患者发音质量,首次解析模型内部机制。
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
- 基于Wav2Vec2的语音识别模型用于自动评估发音清晰度与严重程度。
- 通过层分析发现关键特征层,验证不同预训练数据的影响。
- 结合CCA与可视化提升模型可解释性,助力临床应用。
随着自监督学习(SSL)和语音识别(ASR)技术的发展,基于Wav2Vec2的ASR模型已被微调用于自动化语音障碍质量评估任务,在头颈癌患者语音场景中表现优异,建立了新基准。这表明Wav2Vec2的语音识别维度与临床评估维度高度对齐。然而该系统仍为黑箱,缺乏对模型识别维度与临床评分间关联的清晰解释。本文首次对这一基准模型进行分析,聚焦于发音清晰度与严重程度评估任务。通过逐层分析识别关键特征层,并对比不同预训练数据下的SSL与ASR版本的Wav2Vec2模型性能。此外,采用事后可解释性方法(如典型相关分析CCA与可视化技术),追踪模型演化过程并可视化嵌入表示,显著增强模型可解释性。
原文摘要 · Abstract (English)
With the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pre-trained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。