arXiv:2604.14506cs.CV2026-04中稿 · MIDL 2025

用噪声教师引导的注意力掩码提升医学图像自监督学习效果

Co-distilled attention guided masked image modeling with noisy teacher for self-supervised learning on medical images

论文配图:Co-distilled attention guided masked image modeling with noisy teacher for self-supervised learning on medical images
图 1 · 摘自论文原文
  • 通过注意力引导选择性掩码,减少相邻图像块的信息泄露
  • 引入噪声教师保持注意力头多样性,提升下游任务性能
  • 适用于肺结节分类、肿瘤分割等多类医学图像任务

掩码图像建模(MIM)是一种高效的自监督学习方法,可用于从未标注数据中提取有用特征。然而,传统随机掩码在医学图像上因邻近块语义相似而引发信息泄露,导致自监督学习简化。基于层次化移位窗口(Swin)Transformer虽在医学图像中表现优异,却因缺乏全局[CLS]标记难以应用先进掩码策略。为此,本文提出在协同蒸馏框架内引入注意力引导掩码机制(DAGMaN),有选择地掩码语义共现且具有判别性的图像块,以降低信息泄露并提高预训练难度。但注意力引导掩码会削弱注意力头多样性,影响下游性能。为此,首次将噪声教师引入协同蒸馏框架,在实现注意力掩码的同时保留高注意力头多样性。实验验证了DAGMaN在全样本与少样本肺结节分类、免疫治疗结果预测、肿瘤分割及无监督器官聚类等多个任务上的有效性。

原文摘要 · Abstract (English)

Masked image modeling (MIM) is a highly effective self-supervised learning (SSL) approach to extract useful feature representations from unannotated data. Predominantly used random masking methods make SSL less effective for medical images due to the contextual similarity of neighboring patches, leading to information leakage and SSL simplification. Hierarchical shifted window (Swin) transformer, a highly effective approach for medical images cannot use advanced masking methods as it lacks a global [CLS] token. Hence, we introduced an attention guided masking mechanism for Swin within a co-distillation learning framework to selectively mask semantically co-occurring and discriminative patches, to reduce information leakage and increase the difficulty of SSL pretraining. However, attention guided masking inevitably reduces the diversity of attention heads, which negatively impacts downstream task performance. To address this, we for the first time, integrate a noisy teacher into the co-distillation framework (termed DAGMaN) that performs attentive masking while preserving high attention head diversity. We demonstrate the capability of DAGMaN on multiple tasks including full- and few-shot lung nodule classification, immunotherapy outcome prediction, tumor segmentation, and unsupervised organs clustering.

自监督学习医学图像注意力机制蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。