arXiv:2511.06840cs.CVcs.RO2025-11AAAI被引 9

纯视觉导航新框架,用全景图+动态记忆解决未知环境寻物难题

PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory

  • 基于全景视觉和动态记忆,无需地图与深度传感器
  • 在公开基准上成功率(SR)和路径效率(SPL)均显著领先
  • 适合追求轻量、无地图依赖的机器人导航应用

在未见过的家居环境中实现零样本物体导航(ZSON)仍是家庭机器人面临的挑战,需具备强大的感知与决策能力。现有方法常依赖深度传感器或预建地图,限制了多模态大语言模型(MLLM)的空间推理能力。无地图的ZSON方法虽已出现,但通常只做短视决策,缺乏历史上下文导致陷入局部死锁。本文提出PanoNav,一种全RGB输入、无需地图的ZSON框架,结合全景场景解析模块,挖掘MLLM从全景RGB中获取空间理解的潜力;并通过动态有界记忆队列增强的内存引导决策机制,融合探索历史,避免局部死锁。在公开导航基准上的实验表明,PanoNav在成功率(SR)和路径相似性百分比(SPL)指标上显著优于代表性基线。

原文摘要 · Abstract (English)

Zero-shot object navigation (ZSON) in unseen environments remains a challenging problem for household robots, requiring strong perceptual understanding and decision-making capabilities. While recent methods leverage metric maps and Large Language Models (LLMs), they often depend on depth sensors or prebuilt maps, limiting the spatial reasoning ability of Multimodal Large Language Models (MLLMs). Mapless ZSON approaches have emerged to address this, but they typically make short-sighted decisions, leading to local deadlocks due to a lack of historical context. We propose PanoNav, a fully RGB-only, mapless ZSON framework that integrates a Panoramic Scene Parsing module to unlock the spatial parsing potential of MLLMs from panoramic RGB inputs, and a Memory-guided Decision-Making mechanism enhanced by a Dynamic Bounded Memory Queue to incorporate exploration history and avoid local deadlocks. Experiments on the public navigation benchmark show that PanoNav significantly outperforms representative baselines in both SR and SPL metrics.

零样本导航全景视觉动态记忆机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。