首次为视觉语言模型定义五种基础空间能力并评估短板。
Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from Psychometrics
- 构建心理测量框架,量化五类空间能力:感知、关系、方向、旋转与可视化。
- 主流模型平均得分24.95,远低于人类68.38,3D旋转最弱。
- 小模型如Qwen2-VL-7B表现优于大模型,提示架构限制关键。
多元智能理论强调认知能力的层级性。为推进空间人工智能,我们开创一种心理测量框架,定义视觉语言模型(VLMs)的五种基础空间能力(BSAs):空间感知、空间关系、空间方向、心理旋转和空间可视化。通过九项经验证的心理测量实验,对13个主流VLM进行基准测试,发现其与人类存在显著差距(平均分24.95 vs. 68.38)。关键发现包括:1)VLMs呈现与人类相似的层级结构(二维方向最强,三维旋转最弱),且各能力独立(皮尔逊相关系数r<0.4);2)小型模型如Qwen2-VL-7B表现优于大型模型,其中Qwen得分最高(30.82),InternVL2最低(19.6);3)链式思维(0.100准确率提升)与5样本微调(0.259提升)干预效果受限于架构瓶颈。识别出的主要障碍包括几何编码薄弱与动态模拟缺失。该研究将心理测量学中的BSAs与VLM能力关联,提供空间智能评估诊断工具、具身智能发展方法论基础,并为实现类人空间智能提供认知科学驱动的路线图。
原文摘要 · Abstract (English)
The Theory of Multiple Intelligences underscores the hierarchical nature of cognitive capabilities. To advance Spatial Artificial Intelligence, we pioneer a psychometric framework defining five Basic Spatial Abilities (BSAs) in Visual Language Models (VLMs): Spatial Perception, Spatial Relation, Spatial Orientation, Mental Rotation, and Spatial Visualization. Benchmarking 13 mainstream VLMs through nine validated psychometric experiments reveals significant gaps versus humans (average score 24.95 vs. 68.38), with three key findings: 1) VLMs mirror human hierarchies (strongest in 2D orientation, weakest in 3D rotation) with independent BSAs (Pearson's r<0.4); 2) Smaller models such as Qwen2-VL-7B surpass larger counterparts, with Qwen leading (30.82) and InternVL2 lagging (19.6); 3) Interventions like chain-of-thought (0.100 accuracy gain) and 5-shot training (0.259 improvement) show limits from architectural constraints. Identified barriers include weak geometry encoding and missing dynamic simulation. By linking psychometric BSAs to VLM capabilities, we provide a diagnostic toolkit for spatial intelligence evaluation, methodological foundations for embodied AI development, and a cognitive science-informed roadmap for achieving human-like spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。