arXiv:2509.01360cs.CVcs.LG2025-09被引 2

用自监督学习统一处理多模态医学影像,零样本检索效果领先。

M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

  • 构建混合模态数据集,训练无需定制的统一视觉编码器。
  • 零样本跨模态检索超越基线,未见模态(MRI)也能泛化。
  • 适合追求统一医疗影像模型的研究者与临床开发者。

医学图像检索对临床决策和转化研究至关重要,依赖于判别性视觉表征。然而现有方法在2D、3D和视频医学数据上采用独立架构与训练策略,导致系统碎片化,难以扩展并阻碍统一表征的发展。为此,我们构建了包含867,653个样本的大规模混合模态数据集,涵盖2D X光片与超声、RGB内窥镜视频及3D CT扫描。基于此数据集,我们训练了M3Ret——一个无模态定制的统一视觉编码器,同时采用生成式(MAE)与对比式(SimDINO)自监督学习范式,成功学习可迁移的视觉表示。该方法在所有单模态零样本图像到图像检索任务中达到新基准,优于DINOv3与文本监督的BMC-CLIP等强基线。更显著的是,即使未见过配对数据,仍实现强跨模态对齐,且在未参与预训练的MRI任务上表现出良好泛化能力。全面分析验证了该框架在模型与数据规模上的可扩展性。这些发现为医学影像领域提供重要信号,推动视觉自监督学习迈向多模态医学理解的基础模型。

原文摘要 · Abstract (English)

Medical image retrieval is essential for clinical decision-making and translational research, relying on discriminative visual representations. Yet, current methods remain fragmented, relying on separate architectures and training strategies for 2D, 3D, and video-based medical data. This modality-specific design hampers scalability and inhibits the development of unified representations. To enable unified learning, we curate a large-scale hybrid-modality dataset comprising 867,653 medical imaging samples, including 2D X-rays and ultrasounds, RGB endoscopy videos, and 3D CT scans. Leveraging this dataset, we train M3Ret, a unified visual encoder without any modality-specific customization. It successfully learns transferable representations using both generative (MAE) and contrastive (SimDINO) self-supervised learning (SSL) paradigms. Our approach sets a new state-of-the-art in zero-shot image-to-image retrieval across all individual modalities, surpassing strong baselines such as DINOv3 and the text-supervised BMC-CLIP. More remarkably, strong cross-modal alignment emerges without paired data, and the model generalizes to unseen MRI tasks, despite never observing MRI during pretraining, demonstrating the generalizability of purely visual self-supervision to unseen modalities. Comprehensive analyses further validate the scalability of our framework across model and data sizes. These findings deliver a promising signal to the medical imaging community, positioning M3Ret as a step toward foundation models for visual SSL in multimodal medical image understanding.

医学影像自监督学习多模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。