用手机点云构建可导航的语义空间地图,让AI懂环境、会指路。
GIST: Multimodal Knowledge Extraction and Spatial Grounding via Intelligent Semantic Topology

- 将手机扫描点云转为带语义的2D拓扑图,智能选关键帧与语义
- 一拍即用的导航系统,5米内定位误差仅1.04米,90%成功率
- 适合智能导览、无障碍设计,特别适合复杂室内场景
在零售店、仓库、医院等密集环境中,传统视觉方法因物品静态分布和长尾语义而失效。我们提出GIST(基于智能语义拓扑的空间锚定),将消费级手机点云转化为带有语义标注的导航拓扑结构。系统先生成2D占据图,提取拓扑布局,再通过智能关键帧与语义选择叠加轻量语义层。验证了该结构化空间知识在多项人机交互任务中的有效性:(1) 意图驱动的语义搜索,支持类别替代与区域推断;(2) 一次学习语义定位,5维平均平移误差达1.04米;(3) 区域分类模块,将可行走平面划分为高层语义区;(4) 视觉锚定指令生成器,将路径转化为以地标为中心的自然语言指引。多标准大模型评估中,优于序列式指令生成基线。现场形成性评估(N=5)显示,仅靠口头指令即可实现80%导航成功率,验证其普适设计潜力。
原文摘要 · Abstract (English)
Navigating complex, densely packed environments like retail stores, warehouses, and hospitals poses a significant spatial grounding challenge for humans and embodied AI. In these spaces, dense visual features quickly become stale given the quasi-static nature of items, and long-tail semantic distributions challenge traditional computer vision. While Vision-Language Models (VLMs) help assistive systems navigate semantically-rich spaces, they still struggle with spatial grounding in cluttered environments. We present GIST (Grounded Intelligent Semantic Topology), a multimodal knowledge extraction pipeline that transforms a consumer-grade mobile point cloud into a semantically annotated navigation topology. Our architecture distills the scene into a 2D occupancy map, extracts its topological layout, and overlays a lightweight semantic layer via intelligent keyframe and semantic selection. We demonstrate the versatility of this structured spatial knowledge through critical downstream Human-AI interaction tasks: (1) an intent-driven Semantic Search engine that actively infers categorical alternatives and zones when exact matches fail; (2) a one-shot Semantic Localizer achieving a 1.04 m top-5 mean translation error; (3) a Zone Classification module that segments the walkable floor plan into high-level semantic regions; and (4) a Visually-Grounded Instruction Generator that synthesizes optimal paths into egocentric, landmark-rich natural language routing. In multi-criteria LLM evaluations, GIST outperforms sequence-based instruction generation baselines. Finally, an in-situ formative evaluation (N=5) yields an 80% navigation success rate relying solely on verbal cues, validating the system's capacity for universal design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。