arXiv:2503.04478cs.CV2025-03被引 1

用语义对齐让通用模型学会医学影像理解,零样本提升诊断能力

Semantic Alignment of Unimodal Medical Text and Vision Representations

  • 通过锚点样本估算跨模态变换,实现文本与图像表征对齐
  • 无需额外训练,在多个胸片数据集上显著提升通用模型性能
  • 零样本分类新方法,逼近专用于医疗的多模态模型效果

通用人工智能模型在文本和视觉任务中表现优异,但在医学影像等专业领域常表现不佳,通常需依赖领域特定解决方案或知识迁移。近期研究发现,当处理语义相关数据时,通用模型的潜在空间可具相似性,但这种对齐并非自然发生。基于此,研究显示仅需对少量语义对应样本(锚点)估计一个至多仿射的变换,即可实现跨不同训练范式、架构和模态的模型拼接。本文探索如何利用语义对齐——即通过锚点估计变换——将通用AI模型与医学专业知识相连接。使用多个公开胸片数据集,我们证明跨架构模型拼接可使通用模型在不进行额外训练的情况下融入领域知识,从而提升医学任务表现。此外,我们提出一种新型零样本分类方法,适用于单模态视觉编码器,利用跨模态语义对齐。结果表明,该方法不仅优于通用多模态模型,还接近完全训练的医学专用多模态模型的性能。

原文摘要 · Abstract (English)

General-purpose AI models, particularly those designed for text and vision, demonstrate impressive versatility across a wide range of deep-learning tasks. However, they often underperform in specialised domains like medical imaging, where domain-specific solutions or alternative knowledge transfer approaches are typically required. Recent studies have noted that general-purpose models can exhibit similar latent spaces when processing semantically related data, although this alignment does not occur naturally. Building on this insight, it has been shown that applying a simple transformation - at most affine - estimated from a subset of semantically corresponding samples, known as anchors, enables model stitching across diverse training paradigms, architectures, and modalities. In this paper, we explore how semantic alignment - estimating transformations between anchors - can bridge general-purpose AI with specialised medical knowledge. Using multiple public chest X-ray datasets, we demonstrate that model stitching across model architectures allows general models to integrate domain-specific knowledge without additional training, leading to improved performance on medical tasks. Furthermore, we introduce a novel zero-shot classification approach for unimodal vision encoders that leverages semantic alignment across modalities. Our results show that our method not only outperforms general multimodal models but also approaches the performance levels of fully trained, medical-specific multimodal solutions

医学影像语义对齐零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。