arXiv:2603.13800cs.CV2026-03被引 2

构建首个3D医学影像空间智能评测基准,揭示现有模型空间推理能力不足。

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

  • 通过多智能体协作与放射科医生验证,自动生成3D空间问答数据。
  • 创建包含31253个问题对的SpatialMed基准,覆盖多器官与肿瘤类型。
  • 发现24个顶尖模型在3D医学影像空间推理上普遍表现薄弱。

视觉空间智能对医学图像解读至关重要,但在3D成像的多模态大语言模型(MLLMs)中仍鲜有研究。这一空白源于缺乏结构化3D空间标注数据集。本研究提出一种代理式流水线,通过体积估计、边界框提取等计算工具,结合多智能体协作与专家放射科医生验证,自主合成空间视觉问答(VQA)数据。我们构建了SpatialMed——首个评估医疗MLLMs 3D空间智能的综合性基准,包含跨多个器官和肿瘤类型的31,253个问题-答案对。对24个先进MLLMs的评估及深入分析表明,当前模型在医学影像空间推理方面缺乏稳健能力。

原文摘要 · Abstract (English)

Visual spatial intelligence is critical for medical image interpretation, yet remains largely unexplored in Multimodal Large Language Models (MLLMs) for 3D imaging. This gap persists due to a systemic lack of datasets featuring structured 3D spatial annotations beyond basic labels. In this study, we introduce an agentic pipeline that autonomously synthesizes spatial visual question-answering (VQA) data by orchestrating computational tools such as volume estimation and bounding boxes extraction with multi-agent collaboration and expert radiologist validation. We present SpatialMed, the first comprehensive benchmark for evaluating 3D spatial intelligence in medical MLLMs, comprising 31,253 question-answer pairs across multiple organs and tumor types. Our evaluations on 24 state-of-the-art MLLMs and extensive analyses reveal that current models lack robust spatial reasoning capabilities for medical imaging.

医学影像空间推理多模态模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。