用分层地图与3D高斯表示,让机器人在未知场景中听懂指令并自主探索。
HELIOS: Hierarchical Exploration for Language-Grounded Interaction in Open Scenes
- 构建2D语义地图与3D高斯物体表示,融合多视角观测。
- 在OVMM基准上达最优,真实办公室中成功完成语言指定取放任务。
- 无需任务微调,借助预训练视觉语言模型即可跨仿真与真实场景部署。
在新环境中执行语言指定的移动操作任务时,机器人需应对部分可观测场景、将语言语义对齐至局部感知、以及通过新观测主动更新环境知识等挑战。为此,我们提出HELIOS,一种分层场景表征及相应的搜索目标。该方法构建包含语义与占用信息的2D地图,同时主动构建任务相关物体的3D高斯表示,并通过狄利克雷分布显式建模各物体检测结果的多视角一致性。规划被建模为在分层表征上的搜索问题,目标函数联合考虑(i)未观测或不确定区域的探索,以及(ii)候选物体额外观测带来的信息增益。该目标结合了基于前缘的探索与提升物体检测语义一致性的期望信息增益。我们在Habitat模拟器的OVMM基准上评估HELIOS,该基准为复杂大场景中目标物较小的拾取放置任务,感知难度高。HELIOS在该基准上达到当前最优性能,并在真实办公环境中于Spot机器人上成功演示语言指定的拾取放置任务。本方法利用预训练视觉语言模型,在仿真与真实世界中均无需任务特定训练即取得成果。
原文摘要 · Abstract (English)
Language-specified mobile manipulation tasks in novel environments simultaneously face challenges interacting with a scene which is only partially observed, grounding semantic information from language instructions to the partially observed scene, and actively updating knowledge of the scene with new observations. To address these challenges, we propose HELIOS, a hierarchical scene representation and associated search objective. We construct 2D maps containing the relevant semantic and occupancy information for navigation while simultaneously actively constructing 3D Gaussian representations of task-relevant objects. We fuse observations across this multi-layered representation while explicitly modeling the multi-view consistency of the detections of each object using the Dirichlet distribution. Planning is formulated as a search problem over our hierarchical representation. We formulate an objective that jointly considers (i) exploration of unobserved or uncertain regions of the environment and (ii) information gathering from additional observations of candidate objects. This objective integrates frontier-based exploration with the expected information gain associated with improving semantic consistency of object detections. We evaluate HELIOS on the OVMM benchmark in the Habitat simulator, a pick and place benchmark in which perception is challenging due to large and complex scenes with comparatively small target objects. HELIOS achieves state-of-the-art results on OVMM. We demonstrate HELIOS performing language specified pick and place in a real world office environment on a Spot robot. Our method leverages pretrained VLMs to achieve these results in simulation and the real world without any task specific training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。