arXiv:2606.22694cs.CVcs.SC2026-06

让AI理解多视角空间关系,提升复杂场景推理能力

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

论文配图:SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding
图 1 · 摘自论文原文
  • 构建3D场景重建与视角感知的符号化空间谓词
  • 在3D FORCE和MindCube上分别实现稳定高精度表现
  • 适合需要精准空间推理的视觉语言任务应用

视觉语言模型在依赖参考系的空间关系组合推理中仍不可靠。现有神经符号方法虽使推理更明确,但常依赖脆弱的几何计算和对噪声感知的硬决策。我们提出SATURN,一种面向视角感知的组合式空间推理神经符号框架。SATURN重建近似3D场景,生成软性视角感知的空间谓词,并通过无需训练的Python符号执行器进行组合,分离感知与推理,同时通过多跳推理保留不确定性。我们还引入3D FORCE,一个诊断基准,可控制空间排列接地(SAG)和指代表达接地(REF)中的推理深度、视角与视角组合。在3D FORCE上,VLMs与空间训练模型随深度和视角复杂度增加而性能急剧下降,而SATURN保持稳定并优于强基线。在真实世界MindCube基准上,SATURN整体准确率达78.57%,超越最强基线14个百分点。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing neuro-symbolic methods make reasoning more explicit, but often depend on brittle geometric procedures and hard decisions over noisy perception. We propose SATURN, a neuro-symbolic framework for perspective-aware compositional spatial reasoning. SATURN reconstructs an approximate 3D scene, derives soft perspective-aware spatial predicates, and composes them with a training-free Pythonic symbolic executor, separating perception from reasoning while preserving uncertainty through multi-hop inference. We also introduce 3D FORCE, a diagnostic benchmark that controls reasoning depth, view, and perspective composition across spatial arrangement grounding (SAG) and referring expression grounding (REF). On 3D FORCE, VLMs and spatially trained models degrade sharply as depth and perspective complexity increase, whereas SATURN remains stable and outperforms strong baselines. On the real-world MindCube benchmark, SATURN achieves 78.57% overall accuracy, outperforming the strongest baseline by 14 pp.

空间推理视觉语言神经符号3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。