用网格结构显式建模物体布局,提升大模型的空间智能评估与训练效果。
Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- 提出SIG网格框架,显式编码物体位置与物理先验关系
- 在1.4K驾驶场景帧上验证,空间智能提升显著且稳定
- 适合自动驾驶、视觉推理等需精确空间理解的任务
如何在基础模型中整合与验证空间智能仍是开放挑战。现有方法常以纯文本提示和VQA评分代理视觉-空间智能(VSI),掩盖几何信息,引入语言捷径,削弱对真正空间能力的归因。我们提出空间智能网格(SIG):一种结构化网格表示,显式编码物体布局、对象间关系及物理先验。作为文本的互补通道,SIG为大模型推理提供忠实、可组合的场景结构表征。基于SIG,我们构建了量化模型内在VSI的能力评估指标,有效分离空间能力与语言先验。在少样本上下文学习中,使用GPT与Gemini系列多模态大模型,相比仅用VQA表示,SIG在所有VSI指标上均带来更大、更稳定、更全面的提升,表明其在数据标注与训练中的潜力。我们还发布了SIGBench,一个包含1.4K驾驶帧的基准数据集,附带真实网格标签与人类注视轨迹,支持基于网格的机器空间智能任务及类人注意力驱动的视觉推理任务。
原文摘要 · Abstract (English)
How to integrate and verify spatial intelligence in foundation models remains an open challenge. Current practice often proxies Visual-Spatial Intelligence (VSI) with purely textual prompts and VQA-style scoring, which obscures geometry, invites linguistic shortcuts, and weakens attribution to genuinely spatial skills. We introduce Spatial Intelligence Grid (SIG): a structured, grid-based schema that explicitly encodes object layouts, inter-object relations, and physically grounded priors. As a complementary channel to text, SIG provides a faithful, compositional representation of scene structure for foundation-model reasoning. Building on SIG, we derive SIG-informed evaluation metrics that quantify a model's intrinsic VSI, which separates spatial capability from language priors. In few-shot in-context learning with state-of-the-art multimodal LLMs (e.g. GPT- and Gemini-family models), SIG yields consistently larger, more stable, and more comprehensive gains across all VSI metrics compared to VQA-only representations, indicating its promise as a data-labeling and training schema for learning VSI. We also release SIGBench, a benchmark of 1.4K driving frames annotated with ground-truth SIG labels and human gaze traces, supporting both grid-based machine VSI tasks and attention-driven, human-like VSI tasks in autonomous-driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。