arXiv:2410.06468cs.AIcs.CV2024-10被引 74

测试大模型空间认知能力,发现仍远低于动物水平。

Does Spatial Cognition Emerge in Frontier Models?

  • 构建多尺度空间认知评测基准SPACE,涵盖地图构建、物体布局推理等
  • 多数前沿模型在经典动物认知测试中表现接近随机水平
  • 支持文本与图像双模态评测,适合评估语言与多模态模型

当前尚未。我们提出SPACE基准,系统评估前沿模型的空间认知能力。该基准基于数十年认知科学研究,涵盖生物体在物理环境中移动时所需的大规模地图构建能力、小尺度物体形状与布局推理,以及空间注意力和记忆等认知基础架构。针对多项任务,我们通过文本与图像并行呈现进行测试,可同时评估大型语言模型与大型多模态模型。结果显示,当代前沿模型在多项经典动物认知测试中表现接近随机水平,空间智能仍远未达到动物水平。代码与数据已公开:https://github.com/apple/ml-space-benchmark

原文摘要 · Abstract (English)

Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear when an organism traverses physical environments, smaller-scale reasoning about object shapes and layouts, and cognitive infrastructure such as spatial attention and memory. For many tasks, we instantiate parallel presentations via text and images, allowing us to benchmark both large language models and large multimodal models. Results suggest that contemporary frontier models fall short of the spatial intelligence of animals, performing near chance level on a number of classic tests of animal cognition. Code and data are available: https://github.com/apple/ml-space-benchmark

空间认知大模型评测认知科学多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。