arXiv:2504.05575cs.CVcs.LG2025-04被引 5

轻量级多模态模型高效解答医学图像问题

A Lightweight Large Vision-language Model for Multimodal Medical Images

  • 融合BiomedCLIP与LLaMA-3,专注医学图像理解
  • 80亿参数仅需2张40GB A100显卡,准确率达73.4%
  • 适合临床场景部署,尤其关注开放问答任务

医学视觉问答(VQA)通过解析医学图像并回答临床问题,提升诊疗决策能力。然而,医学影像复杂且模态多样,构建高效高性能的VQA模型极具挑战。本文提出一种轻量级多模态VQA模型,结合BiomedCLIP进行图像特征提取,利用LLaMA-3处理文本。该模型专为医学VQA设计,在OmniMedVQA数据集上达到当前最优性能。模型约含80亿参数,仅需两张40GB A100 GPU即可运行,显著优于更大模型的效率。在开放式问题上准确率达73.4%,超越现有模型,验证其在真实医疗场景中的应用潜力。主要贡献包括专用多模态VQA架构、资源高效设计以及对开放性临床问题的强大回答能力。

原文摘要 · Abstract (English)

Medical Visual Question Answering (VQA) enhances clinical decision-making by enabling systems to interpret medical images and answer clinical queries. However, developing efficient, high-performance VQA models is challenging due to the complexity of medical imagery and diverse modalities. In this paper, we introduce a lightweight, multimodal VQA model integrating BiomedCLIP for image feature extraction and LLaMA-3 for text processing. Designed for medical VQA tasks, our model achieves state-of-the-art performance on the OmniMedVQA dataset. With approximately 8 billion parameters, it requires only two NVIDIA 40 GB A100 GPUs, demonstrating superior efficiency over larger models. Our results show 73.4% accuracy for open-end questions, surpassing existing models and validating its potential for real-world medical applications. Key contributions include a specialized multimodal VQA model, a resource-efficient architecture, and strong performance in answering open-ended clinical questions.

医学VQA轻量模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。