arXiv:2501.13818cs.AIcs.CV2025-01被引 8

用可解释性自动标注医疗模型的虚假关联,提升安全性和可靠性。

Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data

  • 通过可解释性技术实现样本与特征级偏差自动标注
  • 在4个医学数据集上成功识别并缓解模型虚假学习行为
  • 适合医疗AI研发者和临床验证团队使用

深度神经网络在高风险医疗应用中日益普及,但其易受虚假相关性影响而产生捷径学习,可能带来致命后果。现有研究多孤立处理偏差检测或缓解,本文改进了Reveal2Revise框架,引入半自动化可解释性偏差标注能力,支持样本级和特征级标注,为模型消除错误关联提供关键信息。我们在两个模态的四个医学数据集上验证了该框架的有效性,涵盖由数据缺陷引发的可控与真实世界虚假相关性。实验表明,该方法可有效提升VGG16、ResNet50及主流视觉变换器模型的鲁棒性,增强其在真实医疗场景中的适用性。代码已开源:https://github.com/frederikpahde/medical-ai-safety。

原文摘要 · Abstract (English)

Deep neural networks are increasingly employed in high-stakes medical applications, despite their tendency for shortcut learning in the presence of spurious correlations, which can have potentially fatal consequences in practice. Whereas a multitude of works address either the detection or mitigation of such shortcut behavior in isolation, the Reveal2Revise approach provides a comprehensive bias mitigation framework combining these steps. However, effectively addressing these biases often requires substantial labeling efforts from domain experts. In this work, we review the steps of the Reveal2Revise framework and enhance it with semi-automated interpretability-based bias annotation capabilities. This includes methods for the sample- and feature-level bias annotation, providing valuable information for bias mitigation methods to unlearn the undesired shortcut behavior. We show the applicability of the framework using four medical datasets across two modalities, featuring controlled and real-world spurious correlations caused by data artifacts. We successfully identify and mitigate these biases in VGG16, ResNet50, and contemporary Vision Transformer models, ultimately increasing their robustness and applicability for real-world medical tasks. Our code is available at https://github.com/frederikpahde/medical-ai-safety.

医疗AI可解释性偏差检测模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。