arXiv:2511.11034cs.CVcs.AI2025-11

构建医学多模态模型的组合泛化评测基准,测试跨模态、跨解剖、跨任务能力。

CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging

  • 基于模态-解剖-任务三元组框架,统一四类医学数据为问答格式。
  • 模型在未见组合下准确率下降至83.2%,凸显组合泛化的挑战性。
  • 适用于评估零样本、跨任务、模态无关的医学视觉语言模型性能。

近年来,多模态大语言模型实现了视觉与文本输入的统一处理,在通用医疗AI中展现出巨大潜力。然而,其在未见的成像模态、解剖结构和任务类型组合下的组合泛化能力仍待深入探索。我们提出CrossMed,一个基于模态-解剖-任务(MAT)结构的基准,用于评估医疗多模态大模型的组合泛化(CG)。CrossMed将四个公开数据集——CheXpert(X射线分类)、SIIM-ACR(X射线分割)、BraTS 2020(MRI分类与分割)、MosMedData(CT分类)——重构为统一的视觉问答(VQA)格式,生成20,200个多项选择题实例。我们在相关与不相关MAT划分,以及零重叠设置(测试三元组与训练无任何模态、解剖或任务重合)下评估两个开源多模态大模型:LLaVA-Vicuna-7B和Qwen2-VL-7B。在相关划分上,模型分类准确率达83.2%,分割cIoU为0.75;但在不相关与零重叠条件下性能显著下降,证明该基准的难度。我们还发现跨任务迁移效果:仅用分类数据训练也能使分割性能提升7% cIoU。传统模型(ResNet-50、U-Net)仅表现小幅提升,验证了MAT框架的普适性,而多模态大模型在组合泛化方面表现尤为突出。CrossMed为医疗视觉-语言模型的零样本、跨任务、模态无关泛化提供了严格评估平台。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models have enabled unified processing of visual and textual inputs, offering promising applications in general-purpose medical AI. However, their ability to generalize compositionally across unseen combinations of imaging modality, anatomy, and task type remains underexplored. We introduce CrossMed, a benchmark designed to evaluate compositional generalization (CG) in medical multimodal LLMs using a structured Modality-Anatomy-Task (MAT) schema. CrossMed reformulates four public datasets, CheXpert (X-ray classification), SIIM-ACR (X-ray segmentation), BraTS 2020 (MRI classification and segmentation), and MosMedData (CT classification) into a unified visual question answering (VQA) format, resulting in 20,200 multiple-choice QA instances. We evaluate two open-source multimodal LLMs, LLaVA-Vicuna-7B and Qwen2-VL-7B, on both Related and Unrelated MAT splits, as well as a zero-overlap setting where test triplets share no Modality, Anatomy, or Task with the training data. Models trained on Related splits achieve 83.2 percent classification accuracy and 0.75 segmentation cIoU, while performance drops significantly under Unrelated and zero-overlap conditions, demonstrating the benchmark difficulty. We also show cross-task transfer, where segmentation performance improves by 7 percent cIoU even when trained using classification-only data. Traditional models (ResNet-50 and U-Net) show modest gains, confirming the broad utility of the MAT framework, while multimodal LLMs uniquely excel at compositional generalization. CrossMed provides a rigorous testbed for evaluating zero-shot, cross-task, and modality-agnostic generalization in medical vision-language models.

多模态医学影像组合泛化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。