测试3D医学视觉语言模型的空间理解能力,发现其表现远低于预期。
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

- 构建临床真实问题数据集,要求精确解剖定位与三维结构关系推理。
- 8个模型平均仅34%准确率,部分任务低于随机水平。
- 适合关注医疗AI可信性与三维空间推理的研究者。
近年来,3D医学视觉语言模型在联合分析体素图像与文本方面取得进展,展现出在医学视觉问答(VQA)和报告生成中的强大性能。然而,这些模型究竟是基于3D体积学习到空间上精准的解剖知识,还是依赖预训练先验与语言关联仍不明确。这一不确定性源于缺乏对3D医学视觉语言模型语义-空间推理能力的系统评估,而这对临床可靠决策支持至关重要。为此,我们提出CT-SpatialVQA基准,用于评估3D CT数据中的语义-空间推理能力。该基准包含9077个源自1601份放射科报告和对应CT体积的临床真实问答对,通过强化的LLM辅助管道验证,人类共识一致率达95%。数据集要求明确的解剖定位、侧别意识、结构对比及三维结构间关系推理。我们还制定了标准化评估协议,并对8个3D医学视觉语言模型进行测试,发现其在语义-空间推理任务中严重退化,平均准确率仅为34%,部分任务甚至低于随机水平,凸显了将体素证据更深层次融入模型的必要性,以实现可信的临床应用。
原文摘要 · Abstract (English)
Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical visual question-answering (VQA) and report generation. Despite this progress, it remains unclear whether these models learn spatially grounded anatomy from 3D volumes or rely primarily on learned priors and language correlations. This uncertainty stems from the lack of systematic evaluation of semantic-spatial reasoning in volumetric medical VLMs for clinically reliable decision support. To address this gap, we introduce CT-SpatialVQA, a benchmark designed to evaluate semantic-spatial reasoning in 3D CT data. The benchmark comprises 9077 clinically grounded question-answer (QA) pairs derived directly from 1601 radiology reports and CT volumes, which are validated via a robust LLM-assisted pipeline with a 95% human consensus agreement rate. Our dataset requires explicit anatomical localization, laterality awareness, structural comparison, and 3D inter-structure relational reasoning. We also introduce a standardized evaluation protocol and benchmark eight 3D medical VLMs, finding severe degradation on semantic-spatial reasoning tasks, averaging 34% accuracy and often below random, highlighting the need for deeper integration of volumetric evidence for trustworthy clinical use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。