arXiv:2603.17246cs.LG2026-03

调节医学视觉语言模型的模态差距可提升下游性能,且存在最佳中间状态。

On the Cone Effect and Modality Gap in Medical Vision-Language Embeddings

  • 通过单个超参数λ控制模态间隙,无需重训练预训练模型。
  • 缩小过度的模态差距能提升医疗数据上的任务表现,但完全消除非最优。
  • 发现模态间隙是可调属性,适合医学多模态研究者参考。

视觉-语言模型(VLMs)在表示空间中表现出典型的“锥形效应”,即非线性编码器将嵌入映射到高度集中的区域,导致跨模态分离,称为模态差距。尽管这一现象广泛存在,其在监督式多模态学习中的实际影响,特别是在医学领域,仍不明确。本文提出一种轻量级后处理机制,在保持预训练VLM编码器冻结的前提下,通过单一超参数λ持续调控跨模态分离。该方法使我们能在无需昂贵再训练的情况下,系统分析模态差距对下游性能的影响。我们在多种医学和自然图像数据集上评估了通用模型(CLIP、SigLIP)与医学专用模型(BioMedCLIP、MedSigLIP)在监督式多模态设置下的表现。结果一致表明:适度减小模态差距可提升下游性能,且医学数据集对此更敏感;然而,完全消除模态差距并非总是最优,中间程度、任务相关的分离反而效果最佳。这些发现表明,模态差距是多模态表示中可调节的属性,而非应被普遍最小化的量。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) exhibit a characteristic "cone effect" in which nonlinear encoders map embeddings into highly concentrated regions of the representation space, contributing to cross-modal separation known as the modality gap. While this phenomenon has been widely observed, its practical impact on supervised multimodal learning -- particularly in medical domains -- remains unclear. In this work, we introduce a lightweight post-hoc mechanism that keeps pretrained VLM encoders frozen while continuously controlling cross-modal separation through a single hyperparameter {λ}. This enables systematic analysis of how the modality gap affects downstream multimodal performance without expensive retraining. We evaluate generalist (CLIP, SigLIP) and medically specialized (BioMedCLIP, MedSigLIP) models across diverse medical and natural datasets in a supervised multimodal settings. Results consistently show that reducing excessive modality gap improves downstream performance, with medical datasets exhibiting stronger sensitivity to gap modulation; however, fully collapsing the gap is not always optimal, and intermediate, task-dependent separation yields the best results. These findings position the modality gap as a tunable property of multimodal representations rather than a quantity that should be universally minimized.

多模态医学图像表示学习模态差距

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。