arXiv:2511.10946cs.CV2025-11被引 5

用抽象盒子提升视觉语言模型的3D空间理解能力

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

  • 用抽象边界框编码几何与物理关系,构建3D感知管道
  • 零样本测试中在SAT Real上提升8.3%的3D推理准确率
  • 无需额外训练,适合机器人等具身智能应用

视觉语言模型(VLM)在空间认知与物理理解等3D任务上表现不佳,制约了其在机器人和具身智能中的应用。我们发现这是由于3D任务与VLM的2D训练之间存在模态差距,导致从2D输入中低效提取3D信息。为此,提出SandboxVLM框架,通过抽象边界框编码几何结构与物理运动学,构建四阶段3D沙盒重建与感知流程:多视角先验生成、代理高程估计、多视角投票聚类,以及3D感知推理。在多个基准和不同VLM主干网络的零样本设置下评估,该方法持续提升空间智能,在SAT Real上相较基线实现8.3%的性能增益。结果表明,为VLM引入3D抽象可显著增强其3D推理能力,且无需额外训练,为通用具身智能开辟新路径。

原文摘要 · Abstract (English)

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a modality gap between the 3D tasks and the 2D training of VLM, which led to inefficient retrieval of 3D information from 2D input. To bridge this gap, we introduce SandboxVLM, a simple yet effective framework that leverages abstract bounding boxes to encode geometric structure and physical kinematics for VLM. Specifically, we design a 3D Sandbox reconstruction and perception pipeline comprising four stages: generating multi-view priors with abstract control, proxy elevation, multi-view voting and clustering, and 3D-aware reasoning. Evaluated in zero-shot settings across multiple benchmarks and VLM backbones, our approach consistently improves spatial intelligence, achieving an 8.3\% gain on SAT Real compared with baseline methods for instance. These results demonstrate that equipping VLMs with a 3D abstraction substantially enhances their 3D reasoning ability without additional training, suggesting new possibilities for general-purpose embodied intelligence.

3D感知视觉语言模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。