arXiv:2510.04219eess.AScs.SD2025-10被引 6

用探针分析Whisper如何识别口吃语音,提升临床评估可解释性。

Probing Whisper for Dysarthric Speech in Detection and Assessment

  • 通过线性分类器分析编码器各层嵌入,定位关键信息层
  • 中层(13-15层)在病理语音检测中最具判别力
  • 结果对构建可解释的语音障碍评估工具有指导意义

大规模端到端模型如Whisper在多种语音任务中表现优异,但其在病理语音上的内部行为仍不明确。理解口吃语音在不同网络层中的表征,对构建可靠且可解释的临床评估工具至关重要。本研究对Whisper-Medium模型编码器在口吃语音检测与严重程度分类任务中的表现进行探针分析。通过在线性分类器下评估逐层嵌入,并结合轮廓系数和互信息量,从多角度分析层间信息量。为检验模型适应性,还在口吃语音识别任务上对Whisper进行微调后重复分析。结果显示,中层编码器(第13-15层)最具信息量,微调仅带来轻微变化。该研究提升了Whisper嵌入的可解释性,凸显探针分析在指导大模型用于病理语音应用中的价值。

原文摘要 · Abstract (English)

Large-scale end-to-end models such as Whisper have shown strong performance on diverse speech tasks, but their internal behavior on pathological speech remains poorly understood. Understanding how dysarthric speech is represented across layers is critical for building reliable and explainable clinical assessment tools. This study probes the Whisper-Medium model encoder for dysarthric speech for detection and assessment (i.e., severity classification). We evaluate layer-wise embeddings with a linear classifier under both single-task and multi-task settings, and complement these results with Silhouette scores and mutual information to provide perspectives on layer informativeness. To examine adaptability, we repeat the analysis after fine-tuning Whisper on a dysarthric speech recognition task. Across metrics, the mid-level encoder layers (13-15) emerge as most informative, while fine-tuning induces only modest changes. The findings improve the interpretability of Whisper's embeddings and highlight the potential of probing analyses to guide the use of large-scale pretrained models for pathological speech.

语音识别口吃识别模型可解释性大模型探针

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。