用轻量对齐与结构化场景推理,让大模型真正懂空间关系。
SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning
- 通过交叉添加和标记交错实现2D-3D特征轻量对齐,无需大规模预训练。
- 在VSI-Bench上达73.9分,70亿参数超越更大模型表现。
- 适合需要精确空间理解的机器人、自动驾驶等应用。
虽然多模态大语言模型在语义任务中表现出色,但在复杂的几何推理中常缺乏“空间感知”。现有模型通常面临高昂的模态对齐成本和细粒度结构建模精度不足的问题。我们提出SSR框架,通过轻量级对齐机制无缝融合2D与3D表示。为减少训练开销,将3D几何特征锚定于大语言模型已对齐的2D视觉语义,采用跨模态相加与标记交错,有效避免大规模对齐预训练。为支持复杂空间推理,提出新型场景图生成流水线,将全局布局表示为由相对坐标定义的独立局部三元组链。配套增量生成算法,使模型可构建“语言模型友好”的复杂环境结构骨架。进一步拓展至全球尺度3D定位任务,在异构数据源上实现绝对度量精度。在70亿参数规模下,SSR在多个空间智能基准上达到最先进性能,尤其在VSI-Bench上取得73.9分。该方法显著优于更大模型,证明高效特征对齐与结构化场景推理是真实空间智能的核心。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typically suffer from exorbitant modality-alignment costs and deficiency in fine-grained structural modeling precision.We introduce SSR, a framework designed for Structured Scene Reasoning that seamlessly integrates 2D and 3D representations via a lightweight alignment mechanism. To minimize training overhead, our framework anchors 3D geometric features to the large language model's pre-aligned 2D visual semantics through cross-modal addition and token interleaving, effectively obviating the necessity for large-scale alignment pre-training. To underpin complex spatial reasoning, we propose a novel scene graph generation pipeline that represents global layouts as a chain of independent local triplets defined by relative coordinates. This is complemented by an incremental generation algorithm, enabling the model to construct "language-model-friendly" structural scaffolds for complex environments. Furthermore, we extend these capabilities to global-scale 3D global grounding task, achieving absolute metric precision across heterogeneous data sources. At a 7B parameter scale, SSR achieves state-of-the-art performance on multiple spatial intelligence benchmarks, notably scoring 73.9 on VSI-Bench. Our approach significantly outperforms much larger models, demonstrating that efficient feature alignment and structured scene reasoning are the cornerstones of authentic spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。