arXiv:2603.09101cs.CV2026-03被引 1

医学视觉语言模型通过分阶段训练提升诊断能力

MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

  • 按诊断敏感性和样本代表性分两阶段排序训练数据
  • 自适应对比损失使模型在相似病灶间更好区分
  • 在三种影像任务中显著优于现有方法,适合临床应用

医学视觉语言预训练(VLP)模型近期被用于多样化下游任务。然而,现有方法通常同时学习简单与复杂概念,这种非认知的训练方式导致特征表示不佳,尤其在分布外场景下。为此,我们提出知识驱动的认知协调框架(MedKCO),包含预训练数据排序和视觉-语言对比学习目标设计。具体地,基于诊断敏感性和类内样本代表性构建两级课程;针对医学图像类间相似性高问题,引入自适应非对称对比损失,动态调整学习参与度。在三个医学影像场景的多个视觉-语言下游任务上评估,结果表明该方法显著优于多种课程学习基线。

原文摘要 · Abstract (English)

Medical vision-language pretraining (VLP) models have recently been investigated for their generalization to diverse downstream tasks. However, current medical VLP methods typically force the model to learn simple and complex concepts simultaneously. This anti-cognitive process leads to suboptimal feature representations, especially under distribution shift. To address this limitation, we propose a Knowledge-driven Cognitive Orchestration for Medical VLP (MedKCO) that involves both the ordering of the pretraining data and the learning objective of vision-language contrast. Specifically, we design a two level curriculum by incorporating diagnostic sensitivity and intra-class sample representativeness for the ordering of the pretraining data. Moreover, considering the inter-class similarity of medical images, we introduce a self-paced asymmetric contrastive loss to dynamically adjust the participation of the pretraining objective. We evaluate the proposed pretraining method on three medical imaging scenarios in multiple vision-language downstream tasks, and compare it with several curriculum learning methods. Extensive experiments show that our method significantly surpasses all baselines. https://github.com/Mr-Talon/MedKCO.

医学视觉多模态预训练课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。