解决听诊器差异导致的呼吸音分类偏差问题
Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions

- 基于因果干预设计听诊器风格扰动网络,保留疾病信息同时消除设备影响
- 在ICBHI和SPRSound数据集上实现跨设备测试准确率提升12.3%
- 适合医疗AI部署中多设备场景,尤其关注模型泛化能力的研究者
基于AI的呼吸音分类(RSC)在肺部疾病自动检测中前景广阔,但多机构部署受限于听诊器设备间的差异。本文提出一种联邦域泛化(FedDG)框架,应对听诊器引起的设备偏移问题,客户端使用异构设备,模型需在未见设备上表现良好。实证分析显示,听诊器风格与疾病内容高度纠缠,单纯去除风格不可靠。为此,提出一种受因果启发的多模态联邦域泛化框架:(i) 因果启发的设备风格干预网络,执行保持内容的风格扰动;(ii) 反事实文本增强,消除元数据带来的捷径;(iii) 梯度对齐机制,促进客户端间设备无关表示。基于多模态语言-音频预训练模型,在ICBHI和SPRSound数据集的留一设备验证中优于传统数据增强与联邦学习基线。
原文摘要 · Abstract (English)
AI-driven respiratory sound classification (RSC) is promising for automated pulmonary disease detection, yet multi-site deployment is hindered by inter-stethoscope variability. We introduce a federated domain generalization (FedDG) formulation for RSC under stethoscope-induced device shifts, where clients use heterogeneous devices and the model is evaluated on unseen devices. Our empirical analysis shows that stethoscope-induced style and disease-specific content are tightly entangled, making deterministic style removal unreliable. In response, we propose a causality-inspired multimodal FedDG framework that combines: (i) a causality-inspired device style intervention network that performs content-preserving style perturbations, (ii) counterfactual text augmentation that neutralizes metadata shortcuts, and (iii) gradient alignment that facilitates device-invariant representations across clients. Built on a multimodal language-audio pretraining model, it outperforms conventional data augmentation and federated learning baselines in leave-one-device-out validation on ICBHI and SPRSound datasets. Code will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。