arXiv:2505.09372cs.CV2025-05中稿 · ance被引 14

用多维度医学知识增强视觉语言模型,实现零样本皮肤疾病诊断

MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment

  • 将临床描述拆解为多子文本,通过对比学习增强语义对齐
  • 在8个数据集上超越现有模型,零样本分类准确率提升显著
  • 适合医疗AI研究者和需要跨模态理解的皮肤科应用

皮肤科诊断是复杂的多模态挑战,需融合视觉特征与专业临床知识。尽管视觉语言预训练(VLP)推动了医疗AI发展,但其在皮肤科的应用受限于文本长度和缺乏结构化文本。本文提出MAKE框架,一种用于零样本皮肤科任务的多方面知识增强视觉语言预训练方法。针对全面皮肤科描述需涵盖多个知识维度且超出标准文本限制的问题,该框架引入:(1) 多方面对比学习策略,利用大语言模型将临床叙述分解为知识增强的子文本;(2) 细粒度对齐机制,将子标题与诊断相关图像特征关联;(3) 诊断引导的加权方案,根据临床重要性自适应优先处理不同子文本。在403,563对皮肤科图像-文本数据上预训练后,MAKE在八个数据集上的零样本皮肤疾病分类、概念标注和跨模态检索任务中均显著优于当前最优VLP模型。代码将公开于https://github.com/SiyuanYan1/MAKE。

原文摘要 · Abstract (English)

Dermatological diagnosis represents a complex multimodal challenge that requires integrating visual features with specialized clinical knowledge. While vision-language pretraining (VLP) has advanced medical AI, its effectiveness in dermatology is limited by text length constraints and the lack of structured texts. In this paper, we introduce MAKE, a Multi-Aspect Knowledge-Enhanced vision-language pretraining framework for zero-shot dermatological tasks. Recognizing that comprehensive dermatological descriptions require multiple knowledge aspects that exceed standard text constraints, our framework introduces: (1) a multi-aspect contrastive learning strategy that decomposes clinical narratives into knowledge-enhanced sub-texts through large language models, (2) a fine-grained alignment mechanism that connects subcaptions with diagnostically relevant image features, and (3) a diagnosis-guided weighting scheme that adaptively prioritizes different sub-captions based on clinical significance prior. Through pretraining on 403,563 dermatological image-text pairs collected from education resources, MAKE significantly outperforms state-of-the-art VLP models on eight datasets across zero-shot skin disease classification, concept annotation, and cross-modal retrieval tasks. Our code will be made publicly available at https: //github.com/SiyuanYan1/MAKE.

皮肤诊断视觉语言知识增强零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。