arXiv:2503.07487cs.CV2025-03被引 6

让大模型精准识别医学影像,零样本下表现媲美专业模型。

LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?

  • 用解码端特征对齐训练+医学知识锚定模块,提升大模型对医学图像理解。
  • 在零样本疾病识别上超越传统多模态大模型,接近CLIP优化模型性能。
  • 适合医学影像分析、医疗AI研发人员参考,尤其关注零样本泛化能力。

近年来,多模态大语言模型(MLLMs)在视觉-语言任务中展现出卓越的视觉理解与推理能力。然而,我们发现传统视觉问答(VQA)流程中,MLLMs难以有效处理细粒度医学图像数据,因未能充分挖掘模型提取的特征及可用医学知识,导致其在零样本医学疾病识别任务中表现不佳。值得注意的是,这一局限并不意味着MLLMs无法应对细粒度识别问题。从特征表示角度看,它们具备解决此类挑战的巨大潜力。为此,我们提出LLaVA-RadZ,一种利用现有MLLM特征进行零样本医学疾病识别的简单而有效的框架。具体而言,设计了一种端到端训练策略——解码端特征对齐训练(DFAT),以适配MLLM解码器结构,并引入针对不同模态的特定标记;同时提出领域知识锚定模块(DKAM),挖掘大模型内生医学知识,缓解图像-文本对齐中的类别语义鸿沟。大量实验表明,所提方法在零样本疾病识别上显著优于传统MLLMs,性能接近经过高度优化的CLIP基线模型。

原文摘要 · Abstract (English)

Recently, Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in visual understanding and reasoning across various vision-language tasks. However, we found that MLLMs cannot process effectively from fine-grained medical image data in the traditional Visual Question Answering (VQA) pipeline, as they do not exploit the captured features and available medical knowledge fully, results in MLLMs usually performing poorly in zero-shot medical disease recognition. Fortunately, this limitation does not indicate that MLLMs are fundamentally incapable of addressing fine-grained recognition tasks. From a feature representation perspective, MLLMs demonstrate considerable potential for tackling such challenging problems. Thus, to address this challenge, we propose LLaVA-RadZ, a simple yet effective framework for zero-shot medical disease recognition via utilizing the existing MLLM features. Specifically, we design an end-to-end training strategy, termed Decoding-Side Feature Alignment Training (DFAT) to take advantage of the characteristics of the MLLM decoder architecture and incorporate modality-specific tokens tailored for different modalities. Additionally, we introduce a Domain Knowledge Anchoring Module (DKAM) to exploit the intrinsic medical knowledge of large models, which mitigates the category semantic gap in image-text alignment. Extensive experiments demonstrate that our LLaVA-RadZ significantly outperforms traditional MLLMs in zero-shot disease recognition, achieving the comparable performance to the well-established and highly-optimized CLIP-based approaches.

医学影像多模态零样本大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。