构建多模态空间智能评测基准,揭示大模型在空间推理上的短板。
SITE: towards Spatial Intelligence Thorough Evaluation
- 基于认知科学分类体系设计多模态任务,覆盖视点转换与动态场景
- 领先模型在空间定向能力上显著落后于人类专家
- 空间推理能力与具身智能任务表现正相关,具实用参考价值
空间智能(SI)是涵盖空间关系可视化、操作与推理的认知能力,涉及神经科学至机器人学等多个领域。我们提出SITE——一个标准化多选视觉问答格式的基准数据集,旨在全面评估大视觉语言模型在多种视觉模态(单图、多图、视频)和空间智能因素(图形至环境尺度、空间可视化与定向、内在与外在、静态与动态)下的空间智能表现。该基准的构建结合了对31个现有数据集的自下而上调研与自上而下的认知科学分类体系指导,催生出视点转换和动态场景两类新任务。大量实验表明,当前领先模型在空间定向这一核心指标上明显弱于人类专家;此外,我们还证实模型的空间推理能力与其在具身人工智能任务中的表现存在正相关关系。
原文摘要 · Abstract (English)
Spatial intelligence (SI) represents a cognitive ability encompassing the visualization, manipulation, and reasoning about spatial relationships, underpinning disciplines from neuroscience to robotics. We introduce SITE, a benchmark dataset towards SI Thorough Evaluation in a standardized format of multi-choice visual question-answering, designed to assess large vision-language models' spatial intelligence across diverse visual modalities (single-image, multi-image, and video) and SI factors (figural to environmental scales, spatial visualization and orientation, intrinsic and extrinsic, static and dynamic). Our approach to curating the benchmark combines a bottom-up survey about 31 existing datasets and a top-down strategy drawing upon three classification systems in cognitive science, which prompt us to design two novel types of tasks about view-taking and dynamic scenes. Extensive experiments reveal that leading models fall behind human experts especially in spatial orientation, a fundamental SI factor. Moreover, we demonstrate a positive correlation between a model's spatial reasoning proficiency and its performance on an embodied AI task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。