评测大模型的空间想象能力,发现主流模型存在明显短板。
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
- 12个任务分4类,用程序生成1180道题,避免数据污染。
- 27个大模型表现差异大,顶尖模型仍有空间推理缺陷。
- 揭示思维链提示反而降低开源模型表现,值得研究者关注。
人类具备在脑海中想象和操作视觉图像的能力,称为空间可视化。尽管现有多种多模态基准测试关注可见视觉信息的推理,但对通过空间可视化推断不可见关系的能力评估仍不足。依赖公开来源的智商测试或数学竞赛题目存在数据污染风险,影响评估可靠性。为此,我们提出SpatialViz-Bench,一个涵盖4种子能力、共12个任务的综合性多模态基准,包含1,180道程序生成的问题,具有可扩展框架,支持公平且持续可靠的评估。对27个多模态大语言模型(MLLMs)的评估显示性能差异显著,验证了该基准的强大区分能力,并发现反直觉现象:思维链(CoT)提示反而降低开源模型的准确率。通过对错误类型进行统计与定性分析,证明当前先进MLLMs在空间可视化任务中仍存在明显缺陷,填补了领域关键空白。基准数据与评估代码已公开。
原文摘要 · Abstract (English)
Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through spatial visualization remains insufficiently evaluated as a spatial skill. This reliance on publicly sourced problems from IQ tests or math competitions risks data contamination and compromises assessment reliability. To this end, we introduce SpatialViz-Bench, a comprehensive multi-modal benchmark for spatial visualization with 12 tasks across 4 sub-abilities, comprising 1,180 programmatically generated problems, a scalable framework that allows for expansion to ensure fair and continuously reliable evaluations. Our evaluation of 27 Multi-modal Large Language Models (MLLMs) reveals wide performance variations, demonstrates the benchmark's strong discriminative power, and uncovers counter-intuitive findings: Chain-of-Thought (CoT) prompting paradoxically degrades accuracy on open-source models. Through statistical and qualitative analysis of error types, SpatialViz-Bench demonstrates that state-of-the-art MLLMs exhibit deficiencies in spatial visualization tasks, thereby addressing a significant lacuna in the field. The benchmark data and evaluation code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。