arXiv:2512.16461cs.CVcs.RO2025-12被引 4

融合视觉语言模型与三维点云,实现动态环境的时空统一理解。

SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning

  • 用聚类和分割生成物体级提案,结合多模态编码捕捉局部语义、几何与时间特征。
  • 构建可查询的4D场景图,实现空间对齐的跨时序推理,性能达新基准。
  • 无需训练、适配任意主干网络,适合需要时空推理的机器人系统。

自主机器人系统需具备对动态环境的时空理解能力以保障可靠导航与交互。尽管视觉语言模型(VLMs)提供开放世界的语义先验,但缺乏三维几何与时间动态的锚定;而几何感知虽能捕捉结构与运动,却语义稀疏。本文提出SNOW(基于世界知识的场景理解),一种无需训练且主干无关的统一4D场景理解框架,融合VLM语义与点云几何及时间一致性。SNOW处理同步的RGB图像与3D点云,利用HDBSCAN聚类生成物体级提案,指导SAM2分割。每个分割区域通过提出的时空分块编码(STEP)进行编码,生成包含局部语义、几何与时间属性的多模态令牌。这些令牌被增量式整合至4D场景图(4DSG),作为下游推理的4D先验。轻量级SLAM后端将所有STEP令牌在环境中空间锚定,提供全局参考对齐,确保时间上无歧义的空间定位。最终的4DSG形成可查询的统一世界模型,使VLM能直接解析空间结构与时间动态。在多个基准测试中,实验表明SNOW实现了精确的4D场景理解与空间锚定推理,在多个设置中达到新最优性能,凸显结构化4D先验对具身推理与自主机器人的关键作用。

原文摘要 · Abstract (English)

Autonomous robotic systems require spatio-temporal understanding of dynamic environments to ensure reliable navigation and interaction. While Vision-Language Models (VLMs) provide open-world semantic priors, they lack grounding in 3D geometry and temporal dynamics. Conversely, geometric perception captures structure and motion but remains semantically sparse. We propose SNOW (Scene Understanding with Open-World Knowledge), a training-free and backbone-agnostic framework for unified 4D scene understanding that integrates VLM-derived semantics with point cloud geometry and temporal consistency. SNOW processes synchronized RGB images and 3D point clouds, using HDBSCAN clustering to generate object-level proposals that guide SAM2-based segmentation. Each segmented region is encoded through our proposed Spatio-Temporal Tokenized Patch Encoding (STEP), producing multimodal tokens that capture localized semantic, geometric, and temporal attributes. These tokens are incrementally integrated into a 4D Scene Graph (4DSG), which serves as 4D prior for downstream reasoning. A lightweight SLAM backend anchors all STEP tokens spatially in the environment, providing the global reference alignment, and ensuring unambiguous spatial grounding across time. The resulting 4DSG forms a queryable, unified world model through which VLMs can directly interpret spatial scene structure and temporal dynamics. Experiments on a diverse set of benchmarks demonstrate that SNOW enables precise 4D scene understanding and spatially grounded inference, thereby setting new state-of-the-art performance in several settings, highlighting the importance of structured 4D priors for embodied reasoning and autonomous robotics.

4D理解具身推理多模态融合机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。