开源医学多模态模型OpenMedQ在病理视觉问答上表现超越大得多的模型。
OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models
- 在14个公开医疗数据集共335万样本上预训练,覆盖病理、影像、显微及临床问答。
- 路径学视觉问答任务上达到75.9的BLEU-1,超越高达5620亿参数的Med-PaLM M模型。
- 模型可直接迁移至8个新任务,平均宏F1达0.757,适合医学多模态研究者使用。
我们提出OpenMedQ,一个在迄今为止最广泛的全开源医学混合数据上预训练的医学视觉语言模型:涵盖病理学、放射学、显微学和纯文本临床问答的14个数据集,总计约335万条预训练样本。OpenMedQ在PathVQA任务上达到75.9的BLEU-1,显著优于参数量高达562B(约80倍更大)的Med-PaLM M变体,并与最佳报告的VQA-MED BLEU-1(64.5)持平。其视觉编码器在相同下游设置下迁移至8个未见过的医学分类基准,获得最高平均宏F1(0.757),优于BiomedCLIP(0.745)、PMC-CLIP(0.745)、PubMedCLIP(0.746)及从头训练基线(0.616)。代码与交互式演示已公开,为社区提供可复现的基准。
原文摘要 · Abstract (English)
We present OpenMedQ, a medical vision-language model pretrained on the broadest fully-open medical mix to date: 14 datasets totaling ~3.35M pretraining samples spanning pathology, radiology, microscopy, and text-only clinical QA. OpenMedQ reaches state-of-the-art BLEU-1 on PathVQA (75.9), beating Med-PaLM M variants up to 562B parameters (~80x larger), and matches the best reported VQA-MED BLEU-1 (64.5). Its vision encoder, transferred to 8 unseen medical classification benchmarks under an identical downstream recipe, obtains the highest average macro-F1 (0.757) among BiomedCLIP (0.745), PMC-CLIP (0.745), PubMedCLIP (0.746), and a from-scratch baseline (0.616). We release our code and an interactive demo is publicly available as a reproducible baseline for the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。