评测大模型在4维空间推理上的能力,发现普遍不足。
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
- 构建包含4万题的多任务4D空间推理基准
- 模型在路径规划等任务上表现明显落后于人类
- 适合研究多模态模型空间认知能力的学者
4D空间智能涉及对物体随时间运动或变化的感知与处理。人类天然具备4D空间智能,支持广泛的时空推理能力。多模态大模型(MLLMs)在多大程度上能达到人类水平的4D空间智能?本文提出Spatial4D-Bench,一个大规模、多任务的4D空间智能评估基准,涵盖约4万组问答对,覆盖18个明确任务。这些任务被系统划分为六类认知能力:物体理解、场景理解、空间关系理解、时空关系理解、空间推理和时空推理。该基准为评估MLLMs的空间认知能力提供了结构化且全面的测试体系,覆盖了与人类空间智能相当的多样化任务。我们在该基准上评测了多种主流开源及专有MLLMs,揭示其在路径规划、动作识别、物理合理性推理等方面存在显著短板。本研究希望为社区提供关键洞察,推动更接近人类水平的4D空间智能大模型发展。更多资源见项目页面。
原文摘要 · Abstract (English)
4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what extent can Multimodal Large Language Models (MLLMs) achieve human-level 4D spatial intelligence? In this work, we present Spatial4D-Bench, a versatile 4D spatial intelligence benchmark designed to comprehensively assess the 4D spatial reasoning abilities of MLLMs. Unlike existing spatial intelligence benchmarks that are often small-scale or limited in diversity, Spatial4D-Bench provides a large-scale, multi-task evaluation benchmark consisting of ~40,000 question-answer pairs covering 18 well-defined tasks. We systematically organize these tasks into six cognitive categories: object understanding, scene understanding, spatial relationship understanding, spatiotemporal relationship understanding, spatial reasoning and spatiotemporal reasoning. Spatial4D-Bench thereby offers a structured and comprehensive benchmark for evaluating the spatial cognition abilities of MLLMs, covering a broad spectrum of tasks that parallel the versatility of human spatial intelligence. We benchmark various state-of-the-art open-source and proprietary MLLMs on Spatial4D-Bench and reveal their substantial limitations in a wide variety of 4D spatial reasoning aspects, such as route plan, action recognition, and physical plausibility reasoning. We hope that the findings provided in this work offer valuable insights to the community and that our benchmark can facilitate the development of more capable MLLMs toward human-level 4D spatial intelligence. More resources can be found on our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。