arXiv:2410.04521cs.CV2024-10被引 31

用模块化协作思维链提升大模型零样本医学问答能力

MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration

  • 构建模块化协同思维链框架,融合大语言模型与多模态模型各自优势
  • 在SLAKE、VQA-RAD等数据集上准确率超越现有方法,召回率显著提升
  • 适合需要零样本医学图像问答的医疗AI研究者与开发者参考

近期,多模态大语言模型(MLLM)通过在特定医学图像数据集上微调以解决医学视觉问答(Med-VQA)任务。然而,这种针对特定任务的微调方法成本高,且需为每个下游任务单独部署模型,限制了零样本能力的探索。本文提出MC-CoT,一种模块化跨模态协作思维链(CoT)框架,旨在通过整合大语言模型(LLM)提升MLLM在零样本Med-VQA中的表现。MC-CoT通过让LLM提供多种复杂医学推理链,同时由MLLM根据指令提取医学图像的多样化观察结果,实现信息互补。在SLAKE、VQA-RAD和PATH-VQA等数据集上的实验表明,MC-CoT在召回率和准确率上均优于独立的MLLM及多种多模态CoT框架。结果强调了在复杂零样本医学问答任务中引入背景知识与详细引导的重要性。

原文摘要 · Abstract (English)

In recent advancements, multimodal large language models (MLLMs) have been fine-tuned on specific medical image datasets to address medical visual question answering (Med-VQA) tasks. However, this common approach of task-specific fine-tuning is costly and necessitates separate models for each downstream task, limiting the exploration of zero-shot capabilities. In this paper, we introduce MC-CoT, a modular cross-modal collaboration Chain-of-Thought (CoT) framework designed to enhance the zero-shot performance of MLLMs in Med-VQA by leveraging large language models (LLMs). MC-CoT improves reasoning and information extraction by integrating medical knowledge and task-specific guidance, where LLM provides various complex medical reasoning chains and MLLM provides various observations of medical images based on instructions of the LLM. Our experiments on datasets such as SLAKE, VQA-RAD, and PATH-VQA show that MC-CoT surpasses standalone MLLMs and various multimodality CoT frameworks in recall rate and accuracy. These findings highlight the importance of incorporating background information and detailed guidance in addressing complex zero-shot Med-VQA tasks.

医学问答多模态零样本思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。