分离呼吸疾病与说话人特征,提升哮喘和慢阻肺恶化检测的准确性和隐私保护。
Speaker-Disentangled Remote Speech Detection of Asthma and COPD Exacerbations

- 用对抗训练分离语音中的疾病特征与说话人信息,避免身份干扰诊断。
- 在TACTICAS数据集上,疾病状态分类AUC达0.910,类型分类AUC达0.793。
- 通过SHAP分析可解释特征贡献,适合临床辅助诊断与跨数据集应用。
哮喘和慢性阻塞性肺病(COPD)恶化的早期检测对及时干预至关重要。语音作为连续、无创的呼吸疾病监测工具潜力巨大,但语音信号中包含的说话人特征可能主导模型预测,影响诊断性能并威胁患者隐私。此外,呼吸疾病与说话人身份相关的声学特征尚不明确。本文提出一种对抗学习架构,将病理相关声学模式与说话人可识别属性解耦。框架优化两项临床层级任务:(i) 呼吸状态分类(稳定 vs. 恶化),(ii) 恶化类型分类(哮喘恶化 vs. COPD恶化)。通过基于梯度反转的对抗训练抑制说话人身份信息。为增强临床可解释性,采用SHapley Additive exPlanations(SHAP)量化声学特征对病理预测与说话人身份的贡献。在TACTICAS数据集上,该方法在两项任务上均优于单任务基线:呼吸状态分类的AUC从0.897提升至0.910;恶化类型分类的AUC从0.674提升至0.793。同时J-ratio下降,验证了说话人信息的有效抑制。SHAP分析揭示了特征对两类任务的贡献。在Bridge2AI-Voice数据集上的外部验证进一步证实性能提升与说话人依赖性的降低,表明方法具备跨数据集泛化能力。
原文摘要 · Abstract (English)
Early detection of exacerbations in asthma and chronic obstructive pulmonary disease (COPD) is important for timely intervention. Speech has emerged as a promising tool for continuous, non-invasive respiratory disease monitoring. However, speech signals inherently carry speaker-identifiable attributes that may dominate model predictions, which may compromise both diagnosis performance and patient privacy. Furthermore, the acoustic features associated with respiratory disease and speaker identity remain unclear in respiratory disease monitoring. We propose an adversarial learning architecture that disentangles pathology-related acoustic patterns from speaker-identifiable attributes. The framework optimizes two clinically hierarchical tasks: (i) respiratory status classification (stable vs. exacerbated) and (ii) exacerbation type classification (asthma exacerbation vs. COPD exacerbation). Speaker identity is suppressed through gradient reversal-based adversarial training. To enhance clinical interpretability, we employ SHapley Additive exPlanations (SHAP) to quantify the contributions of acoustic features to pathology-related predictions versus speaker identity. On the TACTICAS dataset, our method outperforms the single-task baseline across both tasks. For the respiratory status task (stable vs. exacerbated), the AUC improves from 0.897 to 0.910. For the exacerbation type task (asthma exacerbation vs. COPD exacerbation), the AUC increases from 0.674 to 0.793. Concurrently, the J-ratio decreases, confirming effective suppression of speaker information. SHAP analysis reveals the contributions of the acoustic features to both tasks. External validation on the Bridge2AI-Voice dataset further demonstrates consistent performance improvement and reduced speaker dependency, confirming cross-dataset generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。