arXiv:2606.19776cs.CV2026-06

用2D图像+占用网格实现3D场景理解,让视觉语言模型更懂空间。

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding

论文配图:Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding
图 1 · 摘自论文原文
  • 仅用带位姿的RGB图像和2D视觉编码器,通过重建3D占用网格增强空间感知。
  • 在多视角占用预测上达到顶尖水平,在3D-VQA和密集描述任务上媲美3D输入模型。
  • 适合需要高效3D理解的机器人、具身智能系统,无需复杂3D输入设备。

近期视觉语言模型(VLMs)在3D场景理解方面取得显著进展,推动了具身智能与机器人视觉等应用的发展。然而,现有方法通常依赖显式的3D输入(如点云或RGB-D序列),或引入额外的3D几何编码器,从2D图像中提取3D感知的视觉标记。这类设计在结构上将3D几何感知与基于视觉语言预训练学习到的丰富2D语义解耦,阻碍了统一3D视觉语言表示的发展。本文提出Occ-VLM,一种纯基于带位姿的RGB图像并仅使用单个2D视觉编码器的新框架。具体而言,Occ-VLM通过重建3D场景占用作为辅助几何先验,将前景2D标记与3D空间进行空间关联,再由大语言模型(LLM)解码以实现统一场景理解。大量实验表明,Occ-VLM在准确几何感知与鲁棒视觉语言推理方面均表现优异:在多视角占用预测任务上达到当前最优性能,同时在3D视觉问答(VQA)和3D密集描述基准上表现与依赖3D输入的VLM相当。

原文摘要 · Abstract (English)

Recently, vision-language models (VLMs) have made significant progress in 3D scene understanding, driving advances in applications such as embodied intelligence and robotic vision. However, existing approaches typically either rely directly on explicit 3D inputs (e.g., point clouds or RGB-D sequences), or introduce an additional 3D geometry encoder to derive 3D-aware visual tokens from 2D images. Such designs structurally decouple 3D geometric perception from the rich 2D semantics learned via vision-language pre-training, hindering the development of a unified 3D vision-language representation. In this work, we propose Occ-VLM, a novel framework for 3D scene understanding that operates purely on posed RGB images and employs a single 2D vision encoder. Specifically, Occ-VLM reconstructs 3D scene occupancy as an auxiliary geometric prior, which is utilized to spatially associate foreground 2D tokens with 3D space. These tokens are then decoded by a Large Language Model (LLM) for unified scene understanding. Extensive experiments demonstrate that Occ-VLM achieves both accurate geometric perception and robust vision-language reasoning: it attains state-of-the-art performance on multi-view occupancy prediction, while performing on par with 3D-input VLMs on 3D Visual Question Answering (VQA) and 3D dense captioning benchmarks.

3D理解视觉语言模型占用网格机器人视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。