让机器人看懂物流环境中的物体能动性,无需训练就能理解上下文。
Vision-Language Model Reasoning for Contextual Semantic Mapping in Intralogistics

- 用视觉语言模型多视角推理,零样本推断物体可移动性。
- 语义分类mIoU达98.93%,物体可移动性识别准确率89.17%。
- 适合需要动态环境理解的仓储机器人系统开发者。
在内部物流环境中,自主移动机器人依赖几何地图进行定位与导航,但缺乏对物体及其上下文属性的语义理解。本文提出一种上下文语义映射流程,结合基于SLAM的几何建图、基于SAM的实例分割、实例聚类以及视觉语言模型(VLM)的多视角推理,生成包含几何结构、物体类别和物体可移动性的上下文语义地图。通过跨多视角观测聚合并以零样本、开放词汇方式查询VLM,该流程在不需任务特定训练或预定义物体类别的情况下,推断出物体的上下文属性——此处以可移动性为例。我们评估了三种VLM在两种提示策略下的表现,并对流程各组件进行了分解分析。结果表明,该流程在语义分类上达到98.93%的mIoU,物体可移动性估计准确率达89.17%。组件分析指出,VLM推理是上下文理解的主要瓶颈,而实例聚类是全景性能的主要限制。生成的语义地图支持上下文感知过滤与在动态物流环境中的鲁棒导航。
原文摘要 · Abstract (English)
Autonomous mobile robots operating in intralogistics environments rely on geometric maps for localization and navigation, but lack semantic understanding of objects and their contextual properties. We present a contextual semantic mapping pipeline that combines SLAM-based geometric mapping, SAM-based instance segmentation, instance clustering, and VLM multi-view reasoning to produce a contextual semantic map representation encoding geometric structure, object class, and object movability. By aggregating observations across multiple viewpoints and querying a VLM in a zero-shot, open-vocabulary setting, the pipeline infers contextual object properties--here demonstrated through movability--without requiring task-specific training or predefined object categories. We evaluate three VLMs under two prompting strategies and conduct a component-wise analysis of the pipeline. The proposed pipeline achieves 98.93 % mIoU for semantic classification and 89.17 % mAcc for object movability estimation. Component analysis identifies VLM reasoning as the primary bottleneck for contextual understanding and instance clustering as the main limitation for panoptic performance. The resulting semantic map supports context-aware filtering and robust navigation in dynamic intralogistics environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。