arXiv:2604.15495cs.AIcs.CV2026-04

用手机点云构建可导航的语义空间地图,让AI懂环境、会指路。

GIST: Multimodal Knowledge Extraction and Spatial Grounding via Intelligent Semantic Topology

论文配图:GIST: Multimodal Knowledge Extraction and Spatial Grounding via Intelligent Semantic Topology
图 1 · 摘自论文原文
  • 将手机扫描点云转为带语义的2D拓扑图,智能选关键帧与语义
  • 一拍即用的导航系统,5米内定位误差仅1.04米,90%成功率
  • 适合智能导览、无障碍设计,特别适合复杂室内场景

在零售店、仓库、医院等密集环境中,传统视觉方法因物品静态分布和长尾语义而失效。我们提出GIST(基于智能语义拓扑的空间锚定),将消费级手机点云转化为带有语义标注的导航拓扑结构。系统先生成2D占据图,提取拓扑布局,再通过智能关键帧与语义选择叠加轻量语义层。验证了该结构化空间知识在多项人机交互任务中的有效性:(1) 意图驱动的语义搜索,支持类别替代与区域推断;(2) 一次学习语义定位,5维平均平移误差达1.04米;(3) 区域分类模块,将可行走平面划分为高层语义区;(4) 视觉锚定指令生成器,将路径转化为以地标为中心的自然语言指引。多标准大模型评估中,优于序列式指令生成基线。现场形成性评估(N=5)显示,仅靠口头指令即可实现80%导航成功率,验证其普适设计潜力。

原文摘要 · Abstract (English)

Navigating complex, densely packed environments like retail stores, warehouses, and hospitals poses a significant spatial grounding challenge for humans and embodied AI. In these spaces, dense visual features quickly become stale given the quasi-static nature of items, and long-tail semantic distributions challenge traditional computer vision. While Vision-Language Models (VLMs) help assistive systems navigate semantically-rich spaces, they still struggle with spatial grounding in cluttered environments. We present GIST (Grounded Intelligent Semantic Topology), a multimodal knowledge extraction pipeline that transforms a consumer-grade mobile point cloud into a semantically annotated navigation topology. Our architecture distills the scene into a 2D occupancy map, extracts its topological layout, and overlays a lightweight semantic layer via intelligent keyframe and semantic selection. We demonstrate the versatility of this structured spatial knowledge through critical downstream Human-AI interaction tasks: (1) an intent-driven Semantic Search engine that actively infers categorical alternatives and zones when exact matches fail; (2) a one-shot Semantic Localizer achieving a 1.04 m top-5 mean translation error; (3) a Zone Classification module that segments the walkable floor plan into high-level semantic regions; and (4) a Visually-Grounded Instruction Generator that synthesizes optimal paths into egocentric, landmark-rich natural language routing. In multi-criteria LLM evaluations, GIST outperforms sequence-based instruction generation baselines. Finally, an in-situ formative evaluation (N=5) yields an 80% navigation success rate relying solely on verbal cues, validating the system's capacity for universal design.

空间感知多模态导航语义地图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。