无需标注数据,让机器人在陌生环境找物品更准更可靠
Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation
- 用GLIP+InstructionBLIP双模型检测并验证目标物体
- 真实和仿真环境下导航准确率显著提升
- 适合需要自主探索的机器人应用场景
我们提出一种创新方法,提升零样本物体目标导航中的语义理解能力,增强机器人在陌生环境中的自主性。传统依赖标注数据的方法限制了机器人的适应能力,我们采用由GLIP视觉语言模型与InstructionBLIP模型组成的双组件框架,实现初始检测与语义验证。该方法不仅优化了物体与环境识别,还强化了导航决策所需的语义解释。在模拟与真实场景中进行严格测试,结果表明导航精度与可靠性均有明显提升。
原文摘要 · Abstract (English)
We introduce an innovative approach to advancing semantic understanding in zero-shot object goal navigation (ZS-OGN), enhancing the autonomy of robots in unfamiliar environments. Traditional reliance on labeled data has been a limitation for robotic adaptability, which we address by employing a dual-component framework that integrates a GLIP Vision Language Model for initial detection and an InstructionBLIP model for validation. This combination not only refines object and environmental recognition but also fortifies the semantic interpretation, pivotal for navigational decision-making. Our method, rigorously tested in both simulated and real-world settings, exhibits marked improvements in navigation precision and reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。