arXiv:2506.14512cs.CV2025-06被引 8

构建复杂空间推理数据集,挑战视觉语言模型的结构化空间理解能力

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks

  • 用数学问题自动生成3D场景,构造9000个视频问答对
  • 顶尖视觉语言模型在该基准上表现显著不佳
  • 适合关注空间推理与真实世界交互的AI研究者

大型语言模型的快速进展主要得益于复杂推理任务上的强化学习。相比之下,尽管空间智能对视觉语言模型在现实交互中至关重要,但其复杂空间推理能力的系统性研究仍显不足。为此,我们提出SIRI-Bench,一个通过空间扎根推理任务评估视觉语言模型结构化空间智能的基准。该基准包含9,000个视频-问题-答案三元组,每个问题嵌入于真实的3D场景中,解决需结合空间理解与结构推理。为支持大规模数据合成,我们开发了自动场景生成引擎,利用协作的LLM代理将抽象数学问题转化为忠实的3D场景。实验结果表明,当前最先进视觉语言模型在SIRI-Bench上表现显著不佳,凸显结构化空间推理的挑战。我们希望本研究能引起学界对空间接地推理的关注,推动视觉语言模型在视觉问题求解方面的进步。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, while spatial intelligence is fundamental for Vision-Language Models (VLMs) in real-world interaction, the systematic study of their complex spatial reasoning remains underexplored. To bridge this gap, we introduce SIRI-Bench, a benchmark designed to evaluate VLMs' structural spatial intelligence through spatial-grounded reasoning tasks. SIRI-Bench comprises 9,000 video-question-answer triplets, where each problem is embedded in a realistic 3D scene. The benchmark is carefully designed so that solving each problem requires both spatial comprehension and structural reasoning. To facilitate large-scale data synthesis, we develop an Automatic Scene Creation Engine that employs collaborative LLM agents to translate abstract mathematical problems into faithful 3D scenes. Experimental results reveal that state-of-the-art VLMs struggle significantly on SIRI-Bench, underscoring the challenge of structural spatial reasoning. We hope that our study will bring researchers' attention to spatially grounded reasoning and advance VLMs in visual problem-solving.

空间推理视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。