arXiv:2410.15074cs.CVcs.AI2024-10被引 36

专为中文超声医学设计的多模态大模型,提升医疗图像问答准确率。

LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound

  • 采用细粒度视觉编码融合模块,增强对细微医学图像语义的理解。
  • 在三个医学图像问答数据集上超越现有最优模型,准确率显著提升。
  • 适合医疗AI研究者与临床医生使用,尤其关注中文医学视觉对话场景。

多模态大语言模型(MLLM)正成为研究热点,推动对话式生成AI从单一文本向多模态任务演进,并开始深刻影响医学领域。然而,通用视觉语言模型对医学视觉问答(Med-VQA)的理解能力不足,即使专用医疗模型也常给出模糊且缺乏视觉关联的答案。本文提出一种面向中文医学视觉对话的细粒度自适应视觉语言模型架构,通过参数高效微调实现优化。我们设计了包含细粒度视觉编码器的融合模块,以增强对细微医学视觉语义的捕捉;针对医疗场景中常见的一文多图的数据冗余问题,引入基于知识蒸馏的加权评分机制,自适应筛选与文本描述匹配的图像。模型训练基于医院获取的大规模多模态中文超声数据集,指令数据由专业医生撰写,确保高质量微调。实验表明,所提出的超声医学多模态大模型LLaVA-Ultra在三个Med-VQA数据集上均优于现有最先进模型,在多个指标上表现突出。

原文摘要 · Abstract (English)

Multimodal Large Language Model (MLLM) has recently garnered attention as a prominent research focus. By harnessing powerful LLM, it facilitates a transition of conversational generative AI from unimodal text to performing multimodal tasks. This boom begins to significantly impact medical field. However, general visual language model (VLM) lacks sophisticated comprehension for medical visual question answering (Med-VQA). Even models specifically tailored for medical domain tend to produce vague answers with weak visual relevance. In this paper, we propose a fine-grained adaptive VLM architecture for Chinese medical visual conversations through parameter-efficient tuning. Specifically, we devise a fusion module with fine-grained vision encoders to achieve enhancement for subtle medical visual semantics. Then we note data redundancy common to medical scenes is ignored in most prior works. In cases of a single text paired with multiple figures, we utilize weighted scoring with knowledge distillation to adaptively screen valid images mirroring text descriptions. For execution, we leverage a large-scale multimodal Chinese ultrasound dataset obtained from the hospital. We create instruction-following data based on text from professional doctors, which ensures effective tuning. With enhanced model and quality data, our Large Chinese Language and Vision Assistant for Ultrasound (LLaVA-Ultra) shows strong capability and robustness to medical scenarios. On three Med-VQA datasets, LLaVA-Ultra surpasses previous state-of-the-art models on various metrics.

医学多模态超声分析中文大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。