通过知识锚定的流形迁移,让医学视觉语言模型零样本适配低质影像模态。
K-MaT: Knowledge-Anchored Manifold Transport for Cross-Modal Prompt Learning in Medical Imaging
- 将提示分解并锚定临床文本,用最优传输对齐高低模态提示流形
- 在四个跨模态任务上平均准确率提升至44.1%,乳腺影像任务避免灾难性遗忘
- 无需低质影像训练数据,适合医疗AI快速部署到资源受限场景
基于高阶影像(如CT)训练的大规模生物医学视觉语言模型(VLMs)难以迁移到基层低质模态(如放射摄影),易陷入模态特异性捷径。本文提出K-MaT(知识锚定流形迁移)框架,通过提示学习将决策结构迁移到低质模态,无需低质影像训练数据。K-MaT对提示进行分解,将其锚定于临床文本描述,并利用融合格罗莫夫-沃瑟斯坦最优传输,将低质提示流形对齐至视觉引导的高质量空间。我们在四个跨模态基准上评估了K-MaT,包括皮肤镜、乳腺钼靶到超声、以及CT到胸部X光。K-MaT达到当前最优表现,平均调和均值准确率提升至44.1%(优于BiomedCoOp的42.0%),宏平均F1达36.2%。尤其在具有挑战性的乳腺影像任务中,显著缓解了标准方法(如CoOp)的灾难性遗忘问题(准确率从42.0%降至27.0%),保持跨模态鲁棒性。通过最优传输对齐提示流形,为医学VLM的零样本跨模态部署提供高效路径。
原文摘要 · Abstract (English)
Large-scale biomedical vision-language models (VLMs) adapted on high-end imaging (e.g., CT) often fail to transfer to frontline low-end modalities (e.g., radiography), collapsing into modality-specific shortcuts. We propose K-MaT (Knowledge-Anchored Manifold Transport), a prompt-learning framework that transfers decision structures to low-end modalities without requiring low-end training images. K-MaT factorizes prompts, anchors them to clinical text descriptions, and aligns the low-end prompt manifold to the visually-grounded high-end space using Fused Gromov-Wasserstein optimal transport. We evaluate K-MaT on four cross-modal benchmarks, including dermoscopy, mammography to ultrasound, and CT to chest X-ray. K-MaT achieves state-of-the-art results, improving the average harmonic mean of accuracy to 44.1% (from BiomedCoOp's 42.0%) and macro-F1 to 36.2%. Notably, on the challenging breast imaging task, it mitigates the catastrophic forgetting seen in standard methods like CoOp (which drops to 27.0% accuracy on the low-end), preserving robust performance across modalities. Aligning prompt manifolds via optimal transport provides a highly effective route for the zero-shot cross-modal deployment of medical VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。