arXiv:2412.12661cs.AIcs.CL2024-12NeurIPS被引 10

构建大规模医学多模态数据集,提升医疗AI助手的图文理解与生成能力。

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

  • 构建147万条医学多模态指令数据,覆盖影像、报告生成等任务
  • 在12项医学视觉问答任务上超越Chameleon 26%、GPT-4o 18.3%
  • 提供统一评估套件,助力医疗多模态模型研发

近年来,多模态生成技术的发展为构建统一的生物医学助手开辟了新路径,使其能够分析医学图像、回答复杂问题并生成多模态患者报告。然而,现有数据集存在规模小、任务覆盖有限、来源单一等问题。为此,我们提出MedMax,一个大规模多模态生物医学指令微调数据集,包含147万条实例,涵盖图像-文本交替生成、医学图像描述与生成、视觉对话及报告理解等多种任务,覆盖放射学、组织病理学等多个生物医学领域,基于医学文献和YouTube视频构建。随后,我们在MedMax上微调多模态基础模型,在12项下游生物医学视觉问答任务中实现显著性能提升:相比Chameleon模型提升26%,相较GPT-4o提升18.3%。最后,我们引入一个统一的生物医学任务评估套件,以指导混合模态生物医学AI助手的发展。数据、模型与代码已公开于https://mint-medmax.github.io/。

原文摘要 · Abstract (English)

Recent advancements in mixed-modal generative have opened new avenues for developing unified biomedical assistants capable of analyzing biomedical images, answering complex questions about them, and generating multimodal patient reports. However, existing datasets face challenges such as small sizes, limited coverage of biomedical tasks and domains, and a reliance on narrow sources. To address these gaps, we present MedMax, a large-scale multimodal biomedical instruction-tuning dataset for mixed-modal foundation models. With 1.47 million instances, MedMax encompasses a diverse range of tasks, including interleaved image-text generation, biomedical image captioning and generation, visual chat, and report understanding. These tasks span knowledge across diverse biomedical domains, including radiology and histopathology, grounded in medical papers and YouTube videos. Subsequently, we fine-tune a mixed-modal foundation model on the MedMax dataset, achieving significant performance improvements: a 26% gain over the Chameleon model and an 18.3% improvement over GPT-4o across 12 downstream biomedical visual question-answering tasks. Finally, we introduce a unified evaluation suite for biomedical tasks to guide the development of mixed-modal biomedical AI assistants. The data, model, and code is available at https://mint-medmax.github.io/.

多模态医学AI指令微调数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。