首个支持中英双语的医学多模态模型,能精准定位病灶区域并解释推理过程。
Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks

- 构建区域中心任务与数据集MedRegInstruct,让模型关注图像特定解剖区域
- 在8种医学影像模态上实现跨任务最优性能,报告生成与问答准确率显著提升
- 支持中英文双语交互,提升医生对AI决策过程的理解与信任
现有医学多模态大模型多为全局感知,难以定位具体病灶区域。为此,我们提出区域感知型医学多模态大模型MedRegA,首次实现中英双语下图像级与区域级任务的统一处理。通过构建大规模区域标注数据集MedRegInstruct,并融合多种医学多模态语料训练,模型在8种医学影像模态上均表现优异,在视觉问答、报告生成和图像分类任务中达到领先水平。实验表明,该模型不仅能准确识别多模态医学扫描中的解剖结构,还可解释其推理依据,显著增强可解释性与人机交互能力。
原文摘要 · Abstract (English)
Several medical Multimodal Large Languange Models (MLLMs) have been developed to address tasks involving visual images with textual instructions across various medical modalities, achieving impressive results. Most current medical generalist models are region-agnostic, treating the entire image as a holistic representation. However, they struggle to identify which specific regions they are focusing on when generating a sentence. To mimic the behavior of doctors, who typically begin by reviewing the entire image before concentrating on specific regions for a thorough evaluation, we aim to enhance the capability of medical MLLMs in understanding anatomical regions within entire medical scans. To achieve it, we first formulate Region-Centric tasks and construct a large-scale dataset, MedRegInstruct, to incorporate regional information into training. Combining our collected dataset with other medical multimodal corpora for training, we propose a Region-Aware medical MLLM, MedRegA, which is the first bilingual generalist medical AI system to simultaneously handle image-level and region-level medical vision-language tasks across a broad range of modalities. Our MedRegA not only enables three region-centric tasks, but also achieves the best performance for visual question answering, report generation and medical image classification over 8 modalities, showcasing significant versatility. Experiments demonstrate that our model can not only accomplish powerful performance across various medical vision-language tasks in bilingual settings, but also recognize and detect structures in multimodal medical scans, boosting the interpretability and user interactivity of medical MLLMs. Our project page is https://medrega.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。