arXiv:2609.06629cs.CVcs.RO2026-09

用地图和传感器数据增强视觉模型,让机器人理解真实环境的物理位置关系。

Physico-Geospatial Grounded Scene Interpretation for Mobile Robotics

论文配图:Physico-Geospatial Grounded Scene Interpretation for Mobile Robotics
图 1 · 摘自论文原文
  • 融合地图、传感器与大模型生成带地理定位的语义描述
  • 建筑定位F1分数达0.83,路面类型识别达0.64
  • 适合需要精准空间认知的户外机器人应用

深度学习使机器人能在动态非结构化环境中交互。本文提出一种方法,通过整合预训练视觉语言模型输出,结合开放街图建筑数据、道路信息及感知系统获取的位置、时间与度量信息,利用大模型融合多源数据。在大学校园户外数据集上的初步评估中,建筑定位任务F1得分为0.83,路径表面识别为0.64。结果证明该方案可生成具有物理-地理空间依据的自然语言描述。代码与结果已公开于https://datahub.rz.rptu.de/hstr-csrl-public/publications/physico-geospatial-grounded-scene-interpretation。

原文摘要 · Abstract (English)

Recent advancements in deep learning allow robotic agents to interact with dynamic and unstructured environments. Of special interest is the integration of physico-geospatial world knowledge into such systems, either by using physics-aware machine learning models, knowledge graphs to model relationships or spatio-temporal and logical reasoning. In the present work, we introduce an approach to augment the output of pre-trained, unmodified VLMs used for scene interpretation by integrating semantic descriptions, OpenStreetMap building data and street information with positional, temporal and metric information obtained from our sensory systems, fusing this information using LLMs. We apply this concept to an outdoor recording within a university campus, achieving an F1-Score of 0.83 in the task of grounding buildings and 0.64 for path surface grounding on our pilot evaluation set. The results demonstrate the conceptual capability of the proposed solution to deliver physico-geospatial grounded natural language descriptions. Code and results are available at https://datahub.rz.rptu.de/hstr-csrl-public/publications/physico-geospatial-grounded-scene-interpretation

场景理解地理空间机器人多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。