arXiv:2607.21025cs.RO2026-07

零样本导航新框架,能跨楼层避障并应对动态行人。

ZONDA: Zero-shot Object Navigation with Dynamic Avoidance in Multi-floor Environments

论文配图:ZONDA: Zero-shot Object Navigation with Dynamic Avoidance in Multi-floor Environments
图 1 · 摘自论文原文
  • 基于高度差可行走地图实现跨楼层路径规划,无需特定平台训练。
  • 多视角目标验证降低误检率,提升目标识别准确性。
  • 实时追踪预测行人轨迹,生成前瞻避让行为,适合真实复杂环境。

在物体目标导航任务中,现有方法通常局限于静态、单层环境,忽略了跨楼层拓扑结构和动态行人,限制了实际部署。为此,我们提出ZONDA——一种零样本物体导航与动态避障框架。其核心包含三部分:(i) 基于高度差可行走地图的启发式多楼层规划,支持楼梯通行与跨楼层探索,无需特定平台的训练控制器;(ii) 多视图目标验证,通过视觉-语言模型对多尺度观测进行交叉校验,显著减少误报;(iii) 动态行人避让,显式追踪并预测移动行人,生成前瞻行为。在真实 Direct Drive Tech TITA 双足机器人及大量仿真数据集 HM3D 与 MP3D 上评估,ZONDA 表现显著优于现有基线。此外,在动态基准数据集 HM3D-DYNA 上仍保持鲁棒导航能力。

原文摘要 · Abstract (English)

In Object Goal Navigation task, existing methods are typically restricted to static and single-floor environments, ignoring cross-floor topologies and dynamic pedestrian, which limits their real-world deployment. To address these limitations, we propose ZONDA, a zero-shot object navigation with dynamic avoidance framework. In particular, ZONDA integrates three core components: (i) Heuristic multi-floor planning: from height-difference traversable maps, enables stair traversal and cross-floor exploration without a platform-specific learned controller; (ii) Multi-view target verification: cross-checks multi-scale observations with a vision-language model, significantly reducing false positives; and (iii) Dynamic pedestrian avoidance: explicitly tracks and predicts moving pedestrians to generate anticipatory behaviors. Evaluated on a real Direct Drive Tech TITA biped robot and extensive simulations on HM3D and MP3D, ZONDA achieves significantly improved results. Moreover, ZONDA can maintain robust navigation on the dynamic benchmark HM3D-DYNA compared to the existing baseline.

导航跨楼层动态避障零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。