让医学影像模型更可信:抗攻击能力越强,解释越贴近真实骨折位置。
Secure Diagnostics: Adversarial Robustness Meets Clinical Interpretability
- 用对抗攻击测试模型,发现鲁棒性高的模型解释更符合医生标注的骨折区域。
- 经过微调的模型在对抗样本下仍保持高准确率,且解释与临床解剖区域高度一致。
- 适合关注AI医疗可解释性与安全落地的研究者和临床医生。
用于医学图像分类的深度神经网络常因违反独立同分布假设且决策过程不透明,在临床实践中表现不稳定。本文通过评估微调后用于骨折检测的深度模型在对抗攻击下的表现,并将其解释结果与骨科医生标注的骨折区域进行对比,发现具有鲁棒性的模型生成的解释更贴近临床有意义区域,表明鲁棒性促使模型优先关注解剖学相关特征。研究强调可解释性对促进人机协作的重要性:在‘人类在回路’范式下,临床合理的解释能增强信任、支持错误修正,并避免对高风险决策过度依赖AI。本文将鲁棒性与可解释性视为互补指标,以弥合基准性能与安全、可操作临床部署之间的差距。
原文摘要 · Abstract (English)
Deep neural networks for medical image classification often fail to generalize consistently in clinical practice due to violations of the i.i.d. assumption and opaque decision-making. This paper examines interpretability in deep neural networks fine-tuned for fracture detection by evaluating model performance against adversarial attack and comparing interpretability methods to fracture regions annotated by an orthopedic surgeon. Our findings prove that robust models yield explanations more aligned with clinically meaningful areas, indicating that robustness encourages anatomically relevant feature prioritization. We emphasize the value of interpretability for facilitating human-AI collaboration, in which models serve as assistants under a human-in-the-loop paradigm: clinically plausible explanations foster trust, enable error correction, and discourage reliance on AI for high-stakes decisions. This paper investigates robustness and interpretability as complementary benchmarks for bridging the gap between benchmark performance and safe, actionable clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。