arXiv:2509.17492cs.CVcs.AI2025-09被引 2

通过自监督预训练提升多模态病理图像分类效果,解决标注数据少的问题。

Multimodal Medical Image Classification via Synergistic Learning Pre-training

  • 设计一致性、重建与对齐协同学习的自监督预训练框架
  • 在Kvasir和Kvasirv2数据集上达到当前最优分类性能
  • 适合缺乏标注数据的医疗图像多模态分析场景

多模态病理图像在临床诊断中广泛应用,但基于计算机视觉的多模态辅助诊断面临模态融合挑战,尤其在缺乏专家标注数据的情况下。为应对标签稀缺下的模态融合问题,本文提出一种新型“预训练+微调”框架用于多模态半监督医学图像分类。具体地,设计了一种包含一致性、重建与对齐学习的协同学习预训练机制,将一种模态视为另一种模态的增强样本,实现自监督预训练,显著提升基线模型的特征表达能力。随后,设计多模态融合微调方法:不同编码器分别提取原始模态特征,并引入多模态融合编码器进行特征融合。此外,提出一种分布偏移方法以缓解因标注样本不足导致的预测不确定性与过拟合风险。在公开的胃镜图像数据集Kvasir与Kvasirv2上进行了大量实验,定量与定性结果均表明所提方法优于现有最先进分类方法。代码将在https://github.com/LQH89757/MICS发布。

原文摘要 · Abstract (English)

Multimodal pathological images are usually in clinical diagnosis, but computer vision-based multimodal image-assisted diagnosis faces challenges with modality fusion, especially in the absence of expert-annotated data. To achieve the modality fusion in multimodal images with label scarcity, we propose a novel ``pretraining + fine-tuning" framework for multimodal semi-supervised medical image classification. Specifically, we propose a synergistic learning pretraining framework of consistency, reconstructive, and aligned learning. By treating one modality as an augmented sample of another modality, we implement a self-supervised learning pre-train, enhancing the baseline model's feature representation capability. Then, we design a fine-tuning method for multimodal fusion. During the fine-tuning stage, we set different encoders to extract features from the original modalities and provide a multimodal fusion encoder for fusion modality. In addition, we propose a distribution shift method for multimodal fusion features, which alleviates the prediction uncertainty and overfitting risks caused by the lack of labeled samples. We conduct extensive experiments on the publicly available gastroscopy image datasets Kvasir and Kvasirv2. Quantitative and qualitative results demonstrate that the proposed method outperforms the current state-of-the-art classification methods. The code will be released at: https://github.com/LQH89757/MICS.

多模态医学图像自监督半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。