arXiv:2606.06696cs.CVcs.AI2026-06被引 1

构建最大规模生物医学多模态基准,评测视觉语言模型的感知能力。

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

论文配图:MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
图 1 · 摘自论文原文
  • 设计覆盖35种子模态的多模态生物医学评估体系。
  • 15个开源与前沿模型在跨尺度、跨模态任务中表现参差不齐。
  • 揭示现有高分榜单掩盖了模型在细节感知与泛化上的缺陷。

视觉与语言模型(VLMs)在生物医学影像工作流中潜力巨大,从胸部X光片中的病灶检测到显微镜下的细胞特征分析均有应用前景。实现这一目标需要强大的细粒度视觉感知能力,模型需准确理解图像中细微特征,并在多种生物尺度、临床场景和成像模态下保持鲁棒性。然而,当前基准测试仍存在局限。为此,我们提出大规模多模态生物医学理解(MMBU)基准,是迄今最大的生物医学视觉语言基准,涵盖35个子模态,包含丰富的结构化元数据。该基准包含未标注分类、标注分类及目标检测的开放与封闭版本,支持对模型在不同生物尺度、临床环境和成像模态下的性能进行系统评估。我们评估了15个开源模型和2个前沿模型,发现尽管医学领域适配带来部分提升,但现有基准报告的高精度常掩盖了模型在视觉感知与域泛化方面的不足。

原文摘要 · Abstract (English)

Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual perception. Models need to correctly interpret subtle features in images, and they must do so across diverse biomedical modalities, scales, and contexts. Nevertheless, current benchmarks remain limited. To address these gaps, we introduce the Massive Multimodal Biomedical Understanding (MMBU) benchmark. It is the largest biomedical vision and language benchmark to date, covering 35 submodalities with rich structured metadata. It includes both open and closed versions of ungrounded classification, grounded classification, and object detection, enabling systematic evaluation of model performance across biological scales, clinical settings, and imaging modalities. Evaluating 15 open-weight and 2 frontier VLMs, we find that while medical adaptation provides measurable gains for some models, the high accuracy often reported on established benchmarks can mask deficiencies in visual perception and domain generalization.

多模态生物医学视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。