提出早期干预框架,提升多模态医学图像疾病识别效果
EI: Early Intervention for Multimodal Imaging based Disease Recognition
- 用参考模态的语义令牌干预目标模态的早期特征提取
- 在三个公开数据集上显著优于现有方法,最高提升5.2个点
- 适合缺乏标注数据的医学图像多模态任务应用
基于多模态医学影像的疾病识别面临两大挑战:一是主流的‘单模态嵌入后融合’范式难以充分利用多模态数据间的互补与相关性;二是标注数据稀缺且医学图像与自然图像存在显著域偏移,限制了视觉基础模型(VFMs)在医学图像嵌入中的应用。为此,本文提出一种新型早期干预(EI)框架:以某一模态为目标,其余为参考,利用参考模态的高层语义令牌作为干预令牌,在早期阶段引导目标模态的嵌入过程。同时,提出低秩混合自适应(MoR)方法,采用不同秩的低秩适配器和权重松弛路由机制,实现参数高效的VFMs微调。在视网膜疾病、皮肤病变及膝关节异常分类三个公开数据集上的大量实验验证了该方法的有效性,相比多个基线方法均有显著提升。
原文摘要 · Abstract (English)
Current methods for multimodal medical imaging based disease recognition face two major challenges. First, the prevailing "fusion after unimodal image embedding" paradigm cannot fully leverage the complementary and correlated information in the multimodal data. Second, the scarcity of labeled multimodal medical images, coupled with their significant domain shift from natural images, hinders the use of cutting-edge Vision Foundation Models (VFMs) for medical image embedding. To jointly address the challenges, we propose a novel Early Intervention (EI) framework. Treating one modality as target and the rest as reference, EI harnesses high-level semantic tokens from the reference as intervention tokens to steer the target modality's embedding process at an early stage. Furthermore, we introduce Mixture of Low-varied-Ranks Adaptation (MoR), a parameter-efficient fine-tuning method that employs a set of low-rank adapters with varied ranks and a weight-relaxed router for VFM adaptation. Extensive experiments on three public datasets for retinal disease, skin lesion, and keen anomaly classification verify the effectiveness of the proposed method against a number of competitive baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。