首个支持中英双语的医学多模态大模型,可处理影像与对话。
BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities
- 构建双语医学多模态模型,支持图文交互与多轮对话。
- 在阿拉伯语任务上性能超越现有模型20%以上,英语超9%。
- 适用于临床报告生成、视觉问答等场景,适合医疗AI研究者使用。
我们提出BiMediX2,一个支持阿拉伯语和英语的生物医学专家级大型多模态模型(LMM),可实现基于文本和图像的医疗交互。该模型支持多轮对话,覆盖放射学、CT和组织病理学等多种医学影像模态。为训练模型,我们构建了包含160万样本的BiMed-V双语医疗数据集,涵盖多轮医疗对话、报告生成和视觉问答(VQA)等任务。同时,我们推出了首个阿拉伯语-英语医学多模态模型评估基准BiMed-MBench,经医学专家验证。BiMediX2在多个医学语言模型与多模态模型基准上表现优异,相比其他开源模型达到领先水平:在BiMed-MBench上,英文任务优于现有方法超过9%,阿拉伯语任务提升超过20%;在UPHILL事实准确性测试中,比GPT-4高出约9%;在医学视觉问答、报告生成与摘要任务中也表现出色。模型、指令集与源代码已公开于https://github.com/mbzuai-oryx/BiMediX2。
原文摘要 · Abstract (English)
We introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. It enables multi-turn conversation in Arabic and English and supports diverse medical imaging modalities, including radiology, CT, and histology. To train BiMediX2, we curate BiMed-V, an extensive Arabic-English bilingual healthcare dataset consisting of 1.6M samples of diverse medical interactions. This dataset supports a range of medical Large Language Model (LLM) and Large Multimodal Model (LMM) tasks, including multi-turn medical conversations, report generation, and visual question answering (VQA). We also introduce BiMed-MBench, the first Arabic-English medical LMM evaluation benchmark, verified by medical experts. BiMediX2 demonstrates excellent performance across multiple medical LLM and LMM benchmarks, achieving state-of-the-art results compared to other open-sourced models. On BiMed-MBench, BiMediX2 outperforms existing methods by over 9% in English and more than 20% in Arabic evaluations. Additionally, it surpasses GPT-4 by approximately 9% in UPHILL factual accuracy evaluations and excels in various medical VQA, report generation, and report summarization tasks. Our trained models, instruction set, and source code are available at https://github.com/mbzuai-oryx/BiMediX2
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。