用语音语言模型提取可解释的语音特征,提升心理健康评估的透明度。
An Audio Language Model-Based Voice Concept Bottleneck Framework for Interpretable Health Assessment

- 基于音频语言模型提取离散语音概念分数
- 在抑郁与构音障碍评估中表现优于基线方法
- 适合需要可解释性的医疗语音分析场景
可解释性在临床决策支持中至关重要。概念瓶颈框架通过将输入表示为人类可理解的概念并仅基于这些概念进行预测来提升可解释性。然而,针对语音健康评估的应用研究仍有限。本研究提出一种基于音频语言模型(ALM)的语音概念瓶颈框架,用于可解释的健康评估。该模型在语音质量评估数据集上微调,以增强对语音概念的理解,并作为独立的概念提取器,为轻量级下游分类器生成离散且可解释的评分。离散的概念评分提供直观解释,而轻量级分类器则便于事后可解释性分析。在抑郁和构音障碍评估任务上的结果表明,所提框架能灵活适配不同健康状况的语音概念,且始终优于基于openSMILE和自监督语音模型的基线方法。
原文摘要 · Abstract (English)
Interpretability is critical in clinical decision support. Concept bottleneck frameworks improve it by representing inputs as human-understandable concepts and restricting predictions solely on them. However, research on their use for voice-based health assessment remains limited. In this study, we propose a voice concept bottleneck framework for interpretable health assessment using an audio language model (ALM). The ALM is fine-tuned on a voice quality assessment dataset to enhance its understanding of voice concepts and serves as an independent concept extractor, producing discrete, interpretable scores for a lightweight downstream classifier. The discrete concept scores provide intuitive interpretation, while the lightweight classifier facilitates post-hoc interpretability analyses. Results on depression and dysarthria assessment tasks demonstrate that the proposed framework can flexibly adapt voice concepts to different health conditions and consistently outperforms openSMILE-based and self-supervised speech model-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。