arXiv:2410.23822cs.CVcs.AI2024-10被引 17

轻量微调医疗多模态大模型,精准定位医学图像中的病灶区域。

Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding

  • 仅调整少量参数,适配医学视觉定位任务
  • 在公开数据集上表现优于GPT-4v,实现高精度定位
  • 适合医疗AI研究者和临床辅助诊断系统开发者

多模态大语言模型(MLLM)继承了大语言模型强大的文本理解能力,并将其拓展至多模态场景,在通用多模态任务中表现优异。然而,在医疗领域,高昂的训练成本与对大量医学数据的需求限制了医疗MLLM的发展。此外,由于答案以自由文本形式输出,需生成特定格式输出的任务(如视觉定位)对MLLM构成挑战。目前尚无针对医疗视觉定位的MLLM工作。为此,我们提出一种参数高效微调方法(PFMVG),用于医疗视觉定位任务——根据简短文本描述在医学图像中定位目标区域。我们在公开的医疗视觉定位基准数据集上验证模型性能,结果表明其表现具有竞争力,显著优于GPT-4v。代码将在同行评审后开源。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) inherit the superior text understanding capabilities of LLMs and extend these capabilities to multimodal scenarios. These models achieve excellent results in the general domain of multimodal tasks. However, in the medical domain, the substantial training costs and the requirement for extensive medical data pose challenges to the development of medical MLLMs. Furthermore, due to the free-text form of answers, tasks such as visual grounding that need to produce output in a prescribed form become difficult for MLLMs. So far, there have been no medical MLLMs works in medical visual grounding area. For the medical vision grounding task, which involves identifying locations in medical images based on short text descriptions, we propose Parameter-efficient Fine-tuning medical multimodal large language models for Medcial Visual Grounding (PFMVG). To validate the performance of the model, we evaluate it on a public benchmark dataset for medical visual grounding, where it achieves competitive results, and significantly outperforming GPT-4v. Our code will be open sourced after peer review.

多模态医疗AI视觉定位轻量微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。