让机器人主动理解环境语义,快速精准定位房间与物体。
Active Semantic Perception
- 用多层场景图建模室内环境,支持从房间到物体的多粒度抽象。
- 基于大模型生成未观测区域的合理场景图,提升探索效率。
- 通过信息增益推理路径选择,适合复杂室内导航任务。
本文提出一种主动语义感知方法,利用场景语义指导探索任务。构建了一个紧凑的多层场景图,可表征大型复杂室内环境的多层次抽象信息,如房间、物体、墙壁、窗户等,以及其几何细节。开发基于大语言模型(LLMs)的流程,根据部分观测结果生成未观测区域的合理场景图。设计信息增益计算方法,基于场景图实现高级空间推理:例如,在客厅的两个门中,一个可能通向厨房,另一个通向卧室。在模拟的3D室内公寓及真实世界的Unitree Go 2机器人上进行评估。定性与定量分析表明,本方法能比现有方法更快、更准确地捕捉环境中高阶与低阶语义信息。
原文摘要 · Abstract (English)
We develop an approach for active semantic perception, which refers to using the semantics of the scene for tasks such as exploration. We build a compact, multi-layer scene graph that can represent large, complex indoor environments at various levels of abstraction, e.g., nodes corresponding to rooms, objects, walls, windows etc., as well as fine-grained details of their geometry. We develop a procedure based on large language models (LLMs) to sample new plausible scene graphs of unobserved regions that are consistent with partial observations of the scene. We develop a procedure to compute the information gain of a potential waypoint upon this scene graph to enable sophisticated spatial reasoning: for example, of the two doors that lead out of the living room, one probably leads to the kitchen and the other to the bedroom. We evaluate our approach in realistic 3D indoor apartments in simulation and also on a Unitree Go 2 robot in the real world. Qualitative and quantitative analysis shows that our approach can pin down high-level and low-level semantic information in the environment quickly and more accurately than existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。