arXiv:2608.30467cs.CV2026-08

提出新方法让肺部X光分类模型更关注肺部,而非乱猜。

Beyond Accuracy: Quantifying Pulmonary Attribution in Anatomy-Guided Chest X-Ray Classification Under Domain Shift

论文配图:Beyond Accuracy: Quantifying Pulmonary Attribution in Anatomy-Guided Chest X-Ray Classification Under Domain Shift
图 1 · 摘自论文原文
  • 用双路注意力融合图像与解剖先验,生成软肺部掩码并指导分类
  • 在新冠数据集上准确率达96.15%,肺部关注度提升至70.86%
  • 在跨域测试中仍保持对肺部的高关注,适合医疗可信性研究

深度学习模型在胸部X光分类中表现优异,却未必真正依赖肺部影像内容。本研究将肺部归因包含性作为独立于诊断性能的可靠性指标进行评估。提出DBCASegNet-MGAP多任务解剖引导框架,通过双向跨骨干注意力融合互补特征,预测软肺部掩码,并利用掩码引导自适应全局平均池化(MGAP)将解剖先验直接融入分类。采用解剖局部能量比(ALR)和高强度累积ALR([email protected])量化肺部归因。实验在三个训练种子下进行:四类内部测试使用COVID-19 Radiography Database,零样本外部结核测试采用锁定的深圳-蒙特利尔协议。在新冠数据集上,模型达到加权F1为$0.9615 \pm 0.0015$,宏ROC-AUC为$0.9906 \pm 0.0007$。与传统GAP对比,替换为MGAP后,ALR从$0.3878 \pm 0.0098$提升至$0.7086 \pm 0.0104$,[email protected]从$0.5265 \pm 0.0101$升至$0.9905 \pm 0.0018$,而加权F1基本不变($0.9618 \pm 0.0015$ vs. $0.9615 \pm 0.0015$)。在跨域迁移至蒙特利尔时,ROC-AUC维持在$0.9080 \pm 0.0043$,肺部归因ALR为$0.6466 \pm 0.0081$,但加权F1下降至$0.7528 \pm 0.0080$,ECE上升至$0.1683 \pm 0.0055$。结果表明,诊断判别力、校准性和肺部归因包含性是独立属性,应在内部测试与外部域偏移下联合评估。

原文摘要 · Abstract (English)

Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely on pulmonary image content. This study evaluates pulmonary attribution containment as an anatomy-related reliability property distinct from diagnostic performance. We propose DBCA-SegNet-MGAP, a multi-task anatomy-guided CNN-Transformer framework that combines complementary feature representations through bidirectional cross-backbone attention, predicts a soft lung mask, and incorporates this anatomical prior directly into classification through Mask-Guided Adaptive Global Average Pooling (MGAP). Pulmonary attribution containment is quantified using the Anatomical Local Energy Ratio (ALR) and high-intensity cumulative ALR ([email protected]). Experiments were repeated across three training seeds using the COVID-19 Radiography Database for four-class internal testing and a locked Shenzhen-to-Montgomery protocol for zero-shot external tuberculosis testing. On COVID-19, the proposed model achieved a weighted F1 of $0.9615 \pm 0.0015$ and macro ROC-AUC of $0.9906 \pm 0.0007$. In an architecture-matched dual-bridge comparison, replacing conventional GAP with MGAP increased ALR from $0.3878 \pm 0.0098$ to $0.7086 \pm 0.0104$ and [email protected] from $0.5265 \pm 0.0101$ to $0.9905 \pm 0.0018$, while weighted F1 remained essentially unchanged ($0.9618 \pm 0.0015$ vs. $0.9615 \pm 0.0015$). Under locked external transfer to Montgomery, ROC-AUC remained $0.9080 \pm 0.0043$ and pulmonary ALR remained $0.6466 \pm 0.0081$, whereas weighted F1 decreased to $0.7528 \pm 0.0080$ and ECE increased to $0.1683 \pm 0.0055$. These findings show that diagnostic discrimination, calibration, and pulmonary attribution containment are distinct model properties and support their joint evaluation under internal testing and external domain shift.

医学影像可解释性跨域泛化肺部分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。