机器人用语义感知方法高效探索环境,边找目标边建3D地图。
VISTA: Open-Vocabulary, Task-Relevant Robot Exploration with Online Semantic Gaussian Splatting
- 基于视角与语义覆盖度规划探索路径,兼顾任务相关性与未知区域发现。
- 硬件实验中在复杂环境中成功率提升6倍,静态数据集重建质量更优。
- 适用于无人机和四足机器人,支持开放词汇搜索,开源可复现。
我们提出VISTA(基于视点的图像选择与语义任务感知),一种面向机器人主动探索的方法,旨在规划具有信息量的轨迹以提升对任务完成至关重要的区域的3D地图质量。给定开放词汇搜索指令(如“找一个人”),VISTA使机器人能在搜索目标的同时,实时构建场景的语义3D Gaussian Splatting重建。机器人通过规划递推视野轨迹,优先考虑与查询的语义相似性,并探索环境中的未见区域。为评估轨迹,VISTA引入一种新型高效视点-语义覆盖率指标,量化3D场景中的几何视角多样性与任务相关性。在静态数据集上,该指标在计算速度和重建质量上优于FisherRF与Bayes' Rays等基线方法。在四旋翼硬件实验中,VISTA在挑战性地图中成功率达基线的6倍,而在较简单地图中性能相当。最后,我们验证了VISTA的平台无关性,已部署于四旋翼无人机与Spot四足机器人。论文接受后将开源代码。
原文摘要 · Abstract (English)
We present VISTA (Viewpoint-based Image selection with Semantic Task Awareness), an active exploration method for robots to plan informative trajectories that improve 3D map quality in areas most relevant for task completion. Given an open-vocabulary search instruction (e.g., "find a person"), VISTA enables a robot to explore its environment to search for the object of interest, while simultaneously building a real-time semantic 3D Gaussian Splatting reconstruction of the scene. The robot navigates its environment by planning receding-horizon trajectories that prioritize semantic similarity to the query and exploration of unseen regions of the environment. To evaluate trajectories, VISTA introduces a novel, efficient viewpoint-semantic coverage metric that quantifies both the geometric view diversity and task relevance in the 3D scene. On static datasets, our coverage metric outperforms state-of-the-art baselines, FisherRF and Bayes' Rays, in computation speed and reconstruction quality. In quadrotor hardware experiments, VISTA achieves 6x higher success rates in challenging maps, compared to baseline methods, while matching baseline performance in less challenging maps. Lastly, we show that VISTA is platform-agnostic by deploying it on a quadrotor drone and a Spot quadruped robot. Open-source code will be released upon acceptance of the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。