arXiv:2505.17436cs.AI2025-05被引 14

构建大模型提升生物医学图文理解,支持多任务与零样本学习。

Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning

  • 基于编码器-解码器架构,扩展出BiomedGPT-Large和XLarge模型。
  • 在23个生物医学数据集上微调,覆盖图文问答、图像分类等任务。
  • 支持指令微调与零样本学习,适合科研人员快速部署多模态应用。

为提升生物医学视觉语言模型能力,通过模型规模扩展、微调与指令调优,开发了两种基于编码器-解码器Transformer架构的生物医学视觉语言模型:BiomedGPT-Large和BiomedGPT-XLarge。在6类多模态生物医学任务(包括1个仅图像任务、3个仅文本任务、2个图文任务)共23个基准数据集上进行微调。对比了新模型与先前的BiomedGPT-Base及文献中知名模型的性能。使用大规模多模态生物医学指令数据集对两模型进行指令调优,并评估其零样本学习表现与对齐准确率。

原文摘要 · Abstract (English)

To advance biomedical vison-language model capabilities through scaling up, fine-tuning, and instruction tuning, develop vision-language models with improved performance in handling long text, explore strategies to efficiently adopt vision language models for diverse multi-modal biomedical tasks, and examine the zero-shot learning performance. We developed two biomedical vision language models, BiomedGPT-Large and BiomedGPT-XLarge, based on an encoder-decoder-based transformer architecture. We fine-tuned the two models on 23 benchmark datasets from 6 multi-modal biomedical tasks including one image-only task (image classification), three language-only tasks (text understanding, text summarization and question answering), and two vision-language tasks (visual question answering and image captioning). We compared the developed scaled models with our previous BiomedGPT-Base model and existing prestigious models reported in the literature. We instruction-tuned the two models using a large-scale multi-modal biomedical instruction-tuning dataset and assessed the zero-shot learning performance and alignment accuracy.

视觉语言模型生物医学AI多模态学习指令调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。