提出新型因果注意力机制,提升多标签医学图像识别的准确与可解释性。
Information Bottleneck-based Causal Attention for Multi-label Medical Image Recognition
- 基于信息瓶颈构建因果注意力模型,分离因果、伪相关与噪声因素。
- 在MuReD和Endo数据集上,多项指标优于现有方法,最高提升7.72%。
- 适合需要高可信诊断的医疗视觉任务,尤其关注可解释性研究者。
医学图像多标签分类旨在识别多种疾病,具有重要临床价值。关键在于学习类别特异性特征以实现精准诊断与可解释性。然而,现有方法虽聚焦因果注意力,却因无意中关注无关特征而难以揭示真实病因。为此,本文提出一种新的结构因果模型(SCM),将类别特异性注意力视为因果、伪相关与噪声因素的混合,并设计信息瓶颈驱动的因果注意力(IBCA)机制,用于医学图像多标签分类。具体地,通过学习高斯混合多标签空间注意力,过滤无关信息并捕捉每类特征模式;再引入对比增强因果干预,逐步削弱伪相关注意力,降低噪声干扰。在Endo与MuReD数据集上的定量与消融实验表明,IBCA显著优于所有对比方法:相比次优结果,在MuReD上CR提升6.35%,OR提升7.72%,mAP提升5.02%;在Endo上CR提升1.47%,CF1提升1.65%,mAP提升1.42%。
原文摘要 · Abstract (English)
Multi-label classification (MLC) of medical images aims to identify multiple diseases and holds significant clinical potential. A critical step is to learn class-specific features for accurate diagnosis and improved interpretability effectively. However, current works focus primarily on causal attention to learn class-specific features, yet they struggle to interpret the true cause due to the inadvertent attention to class-irrelevant features. To address this challenge, we propose a new structural causal model (SCM) that treats class-specific attention as a mixture of causal, spurious, and noisy factors, and a novel Information Bottleneck-based Causal Attention (IBCA) that is capable of learning the discriminative class-specific attention for MLC of medical images. Specifically, we propose learning Gaussian mixture multi-label spatial attention to filter out class-irrelevant information and capture each class-specific attention pattern. Then a contrastive enhancement-based causal intervention is proposed to gradually mitigate the spurious attention and reduce noise information by aligning multi-head attention with the Gaussian mixture multi-label spatial. Quantitative and ablation results on Endo and MuReD show that IBCA outperforms all methods. Compared to the second-best results for each metric, IBCA achieves improvements of 6.35\% in CR, 7.72\% in OR, and 5.02\% in mAP for MuReD, 1.47\% in CR, and 1.65\% in CF1, and 1.42\% in mAP for Endo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。