解决医学多模态对齐中的语义鸿沟问题,提升图像与文本的匹配效果。
Closing the gap in multimodal medical representation alignment
- 提出无模态依赖的对齐框架,统一处理医学图像与临床文本。
- 在放射科图像与临床文本间实现更紧密的跨模态对齐,提升检索准确率。
- 适用于医疗AI场景,尤其适合需要精准图文匹配的研究者。
在多模态学习中,CLIP已成为将不同模态映射到共享潜在空间的标准方法,通过拉近语义相似表示、推远不相似表示来实现对齐。然而,基于CLIP的对比损失存在未预期行为,导致潜在空间稀疏且碎片化,这种现象称为模态鸿沟。尽管标准图文对的鸿沟已部分缓解,但在更复杂的医学多模态场景中仍未知且未解决。本文研究医学对齐中的该现象,发现模态鸿沟同样存在于医学领域,并提出一种无模态依赖的框架,有效弥合这一鸿沟,确保语义相关表示无论来源模态均能更好对齐。该方法显著提升放射科图像与临床文本间的跨模态检索与图像描述生成性能。
原文摘要 · Abstract (English)
In multimodal learning, CLIP has emerged as the de-facto approach for mapping different modalities into a shared latent space by bringing semantically similar representations closer while pushing apart dissimilar ones. However, CLIP-based contrastive losses exhibit unintended behaviors that negatively impact true semantic alignment, leading to sparse and fragmented latent spaces. This phenomenon, known as the modality gap, has been partially mitigated for standard text and image pairs but remains unknown and unresolved in more complex multimodal settings, such as the medical domain. In this work, we study this phenomenon in the latter case, revealing that the modality gap is present also in medical alignment, and we propose a modality-agnostic framework that closes this gap, ensuring that semantically related representations are more aligned, regardless of their source modality. Our method enhances alignment between radiology images and clinical text, improving cross-modal retrieval and image captioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。