用可微分的群体结构建模提升医学图像注意力解析力
DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging
- 将度修正混合成员模型作为注意力偏置,实现可微分的群体结构建模
- 在脑、胸、乳腺、眼等多种医学影像上均超越现有方法
- 生成具有解剖意义的注意力图,增强模型可解释性
医学图像中存在器官、组织和病灶等潜在解剖群组,标准视觉变换器(ViT)难以利用。尽管近期工作如SBM-Transformer通过随机二值掩码引入结构,但面临不可微、训练不稳定及复杂社区结构建模能力不足的问题。本文提出DCMM-Transformer,一种新型医学图像分析视觉变换器,将度修正混合成员(DCMM)模型作为自注意力的加性偏置。与依赖乘性掩码和二值采样的方法不同,本方法以全可微且可解释的方式引入群体结构与节点度异质性。在脑、胸、乳腺、眼等多种医学影像数据集上的综合实验表明,该方法性能优越且泛化能力强。此外,学习到的群体结构与结构化注意力调制显著提升了可解释性,生成的注意力图具有解剖学意义且语义连贯。
原文摘要 · Abstract (English)
Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。