用分层图结构让机器人更智能地探索未知环境
ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration

- 构建房间-视角-物体三层视觉图谱,统一环境表示
- 在有限预算下实现粗到细检索,提升探索效率30%以上
- 适合需要自主探索的机器人任务,如家庭导航
在复杂陌生环境中,机器人需通过探索获取完成任务所需信息。我们提出ObsGraph,一种以观察为中心的分层场景图,统一场景表示、检索与探索。该结构将视觉证据组织为房间-视角-物体三层:房间提供粗粒度语义锚点,视角保留物体共视性上下文,物体存储细粒度细节。在此基础上,我们在有限预算下执行粗到细的层级检索,并利用检索结果构建探索候选空间——激活房间级探索、视角细化或前沿探索,从而紧密耦合表示、检索与自适应多尺度探索。在多个具身推理与探索基准上的实验表明,该方法显著提升成功率与效率,凸显结构化场景表示和基于证据缺口的精准信息采集的优势。
原文摘要 · Abstract (English)
Embodied reasoning and exploration are increasingly considered crucial abilities for robots operating in complex and unfamiliar environments. To accomplish tasks in such settings, an agent must identify and acquire the information necessary for the task through exploration. We propose ObsGraph, an observation-centric hierarchical scene graph that unifies scene representation, retrieval, and exploration. It retains visual evidence and organizes it into room-view-object layers: rooms provide coarse semantic anchors, views preserve contextual object covisibility, and objects store fine-grained details. On top of this representation, we perform coarse-to-fine hierarchical retrieval under a bounded budget, and crucially use retrieval outcomes to structure the exploration candidate space--activating room-level exploration, view refinement, or frontier exploration--thereby tightly coupling representation, retrieval, and adaptive multi-scale exploration. Experiments across embodied reasoning and exploration benchmarks demonstrate improved success and efficiency, highlighting the benefits of structured scene representation and more targeted information gathering driven by identified evidence gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。