arXiv:2606.27973cs.CLcs.AI2026-06中稿 · Interspeech 2026

将语音认知障碍检测模型从黑箱变为可解释的临床工具

From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection

论文配图:From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection
图 1 · 摘自论文原文
  • 分四阶段用大模型解析语音特征,生成医学可读报告
  • 在NIA PREPARE数据集上F1达72.11%,与患者认知状态高度吻合
  • 适合临床医生使用,系统易用性评分82/100

基于语音的认知障碍检测为昂贵的生物标志物检测提供了无创、便捷的替代方案,但基于Transformer的模型仍缺乏临床可解释性。本文提出一种多阶段可解释性框架,通过集成SHAP词元归因、理论指导的语言学特征及基于LLaMA-3.1-70B-Instruct的四阶段大模型推理流程,将黑箱预测转化为临床可理解的叙述。该框架基于SpeechCARE-Adaptive Gating Network多模态筛查模型(在NIA PREPARE基准上F1=72.11%),将模型输出映射至词汇丰富度、句法复杂度和语义连贯性等四个认知语言维度。对70个分层英语样本的医生评估显示,结果与患者认知水平高度一致,系统可用性量表得分为82/100,表明其具有良好的临床流程整合潜力。

原文摘要 · Abstract (English)

Speech-based cognitive impairment detection offers a noninvasive, accessible alternative to costly biomarker assays, yet transformer-based models remain clinically uninterpretable. We propose a multi-stage explainability framework that translates black-box transformer predictions into clinically grounded narratives by integrating SHapley Additive exPlanations (SHAP)-based token attribution, theory-informed linguistic features, and a four-stage LLM reasoning pipeline using LLaMA-3.1-70B-Instruct. Built on the SpeechCARE-Adaptive Gating Network multimodal screening model (F1 = 72.11% on the NIA PREPARE benchmark), the framework maps model outputs to four cognitive-linguistic dimensions, including lexical richness, syntactic complexity, and semantic coherence. Physician evaluation on 70 stratified English samples demonstrated strong alignment with patient-level cognitive profiles, and a System Usability Scale score of 82/100 indicated high potential for clinical workflow integration.

语音识别可解释AI认知障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。