arXiv:2602.19308cs.ROcs.CV2026-02被引 2

让机器人在野外长距离搜索目标,既懂语义又保安全。

WildOS: Open-Vocabulary Object Search in the Wild

  • 用视觉基础模型实时评估探索方向的可通行性与目标相似性。
  • 在多种复杂地形中导航效率比纯几何或纯视觉方法提升显著。
  • 适合需要远距离自主探索的野外机器人任务。

在复杂非结构化户外环境中实现自主导航,要求机器人在无先验地图、深度感知有限的情况下进行长距离探索。仅依赖几何前沿的探索策略往往不足,语义推理能力对判断行进方向与安全性至关重要。本文提出WildOS,一个统一的长距离开放词汇目标搜索系统,结合安全的几何探索与语义视觉推理。WildOS构建稀疏导航图以维持空间记忆,利用基于基础模型的视觉模块ExploRFM对图节点进行评分。ExploRFM同步预测可通行性、视觉前沿与图像空间中的对象相似性,支持实时、机载的语义导航任务。由此生成的视觉评分图使机器人能朝语义有意义的方向探索,同时确保几何安全。此外,我们引入基于粒子滤波的粗略定位方法,估计超出机器人即时深度视野的目标候选位置,从而实现向远距离目标的有效规划。在多种非铺装路面与城市环境中的闭环实地实验表明,WildOS实现了鲁棒导航,在效率和自主性上显著优于纯几何与纯视觉基线。结果表明,视觉基础模型有望驱动兼具语义感知与几何约束的开放世界机器人行为。

原文摘要 · Abstract (English)

Autonomous navigation in complex, unstructured outdoor environments requires robots to operate over long ranges without prior maps and limited depth sensing. In such settings, relying solely on geometric frontiers for exploration is often insufficient. In such settings, the ability to reason semantically about where to go and what is safe to traverse is crucial for robust, efficient exploration. This work presents WildOS, a unified system for long-range, open-vocabulary object search that combines safe geometric exploration with semantic visual reasoning. WildOS builds a sparse navigation graph to maintain spatial memory, while utilizing a foundation-model-based vision module, ExploRFM, to score frontier nodes of the graph. ExploRFM simultaneously predicts traversability, visual frontiers, and object similarity in image space, enabling real-time, onboard semantic navigation tasks. The resulting vision-scored graph enables the robot to explore semantically meaningful directions while ensuring geometric safety. Furthermore, we introduce a particle-filter-based method for coarse localization of the open-vocabulary target query, that estimates candidate goal positions beyond the robot's immediate depth horizon, enabling effective planning toward distant goals. Extensive closed-loop field experiments across diverse off-road and urban terrains demonstrate that WildOS enables robust navigation, significantly outperforming purely geometric and purely vision-based baselines in both efficiency and autonomy. Our results highlight the potential of vision foundation models to drive open-world robotic behaviors that are both semantically informed and geometrically grounded. Project Page: https://leggedrobotics.github.io/wildos/

机器人导航视觉基础模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。