无需训练,用3D场景图提升多模态大模型的空间推理能力
GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs

- 通过轻量3D场景图提供几何、布局和视觉属性支持
- 在ScanQA上使CIDEr提升27%,在VSI-Bench上最高提升65%
- 适用于冻结模型,无需微调,适配各类多模态大模型
3D空间推理是理解与操作物理世界的基础,但当前多模态大语言模型(MLLMs)在精确几何测量、视角转换及细粒度外观定位方面仍表现不佳。现有方法通常依赖大规模标注数据微调或附加专用3D编码器,导致需昂贵监督且绑定特定主干网络。我们提出GraFT,一种无需训练的框架,通过紧凑易维护的3D场景图(3DSG)补充缺失的3D结构。该框架提供三项空间推理能力:(1) 借助符号工具实现确定性几何计算,(2) 通过鸟瞰图(BEV)渲染实现非自我中心布局表达,(3) 通过任务相关的自我中心视图实现视觉属性定位。在ScanQA上,GraFT相较同主干基线所有指标均有提升,CIDEr提高27%;在VSI-Bench上,对冻结的MLLMs提升达65%,超越所有专有及通用开源基线,以及多个知名微调空间模型。
原文摘要 · Abstract (English)
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egocentric and allocentric viewpoints, and at grounding fine-grained appearance. The most common remedies fine-tune the model on large-scale curated spatial-reasoning datasets or attach dedicated encoders for 3D geometry, which typically couples the solution to costly supervision and a specific backbone. We instead introduce GraFT, a training-free framework that supplies the missing 3D structure through a compact, easily maintained 3D scene graph (3DSG). From this 3DSG, GraFT provides three spatial reasoning capabilities: (1) deterministic geometry through symbolic tools, (2) allocentric layout through a bird's-eye-view (BEV) rendering, and (3) visual-attribute grounding through task-relevant egocentric frames. On ScanQA, GraFT improves every metric over the same-backbone baseline, raising CIDEr by 27%. On VSI-Bench, GraFT improves frozen MLLMs by up to 65%, surpassing every proprietary and general-purpose open-source baseline, and several prominent fine-tuned spatial models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。