用随机MLP正则化提升DINOv2在医学图像中的可解释性
Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- 引入随机MLP正则化,通过对比学习让特征更语义一致
- 在医疗与自然图像上均保持性能并改善注意力图可读性
- 适合关注模型透明度与跨域泛化的研究者
视觉变换器(ViT)如DINOv2在跨域任务中表现强劲,但常过度利用低信息量的图像块令牌,降低注意力与特征图的可解释性。这一问题在医学影像中尤为突出,因域偏移可能导致性能下降和透明度丧失。本文提出基于对比学习的随机MLP(RMLP)正则化方法,用于微调DINOv2,在医疗与自然图像模态上均能保持或提升下游性能,并生成更具可解释性的注意力图。我们还提供了RMLP的数学分析,揭示其在增强ViT模型表现及深化对比学习理解中的作用。
原文摘要 · Abstract (English)
Vision Transformers (ViTs), such as DINOv2, achieve strong performance across domains but often repurpose low-informative patch tokens in ways that reduce the interpretability of attention and feature maps. This challenge is especially evident in medical imaging, where domain shifts can degrade both performance and transparency. In this paper, we introduce Randomized-MLP (RMLP) regularization, a contrastive learning-based method that encourages more semantically aligned representations. We use RMLPs when fine-tuning DINOv2 to both medical and natural image modalities, showing that it improves or maintains downstream performance while producing more interpretable attention maps. We also provide a mathematical analysis of RMLPs, offering insights into its role in enhancing ViT-based models and advancing our understanding of contrastive learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。