让机器人在远距离模糊可见目标下稳定导航,不依赖训练数据。
EZREAL: Enhancing Zero-Shot Outdoor Robot Navigation toward Distant Targets under Varying Visibility
- 构建分层图像块系统,融合局部语义生成稳定目标区域显著性。
- 在真实户外环境中实现150米外目标检测,可见性变化时82.6%保持正确朝向。
- 轻量闭环设计,适合移动端部署,零样本适配新场景。
大规模室外环境中的零样本目标导航面临诸多挑战,尤其当目标距离遥远导致投影极小且受部分或完全遮挡影响可见性时。本文提出一种统一、轻量的闭环系统,基于对齐的多尺度图像块层次结构。通过分层目标显著性融合,将局部语义对比归纳为稳定的粗粒度区域显著性,提供目标方向并指示可见性。该显著性支持可见性感知的朝向维持,结合关键帧记忆、历史朝向加权融合及临时不可见时的主动搜索。系统避免全图重缩放,实现确定性的自底向上聚合,支持零样本导航,并在移动机器人上高效运行。在仿真与真实室外测试中,系统可检测超过150米外的语义目标,在可见性变化下以82.6%的概率保持正确朝向,相较最先进方法任务成功率提升17.5%,展现出对远距离、间歇可见目标的鲁棒零样本导航能力。
原文摘要 · Abstract (English)
Zero-shot object navigation (ZSON) in large-scale outdoor environments faces many challenges; we specifically address a coupled one: long-range targets that reduce to tiny projections and intermittent visibility due to partial or complete occlusion. We present a unified, lightweight closed-loop system built on an aligned multi-scale image tile hierarchy. Through hierarchical target-saliency fusion, it summarizes localized semantic contrast into a stable coarse-layer regional saliency that provides the target direction and indicates target visibility. This regional saliency supports visibility-aware heading maintenance through keyframe memory, saliency-weighted fusion of historical headings, and active search during temporary invisibility. The system avoids whole-image rescaling, enables deterministic bottom-up aggregation, supports zero-shot navigation, and runs efficiently on a mobile robot. Across simulation and real-world outdoor trials, the system detects semantic targets beyond 150m, maintains a correct heading through visibility changes with 82.6% probability, and improves overall task success by 17.5% compared with the SOTA methods, demonstrating robust ZSON toward distant and intermittently observable targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。