arXiv:2504.09000cs.RO2025-04被引 14

用分层思维链+闭环反馈提升机器人零样本找物能力

CL-CoTNav: Closed-Loop Hierarchical Chain-of-Thought for Zero-Shot Object-Goal Navigation with Vision-Language Models

  • 通过多轮问答数据微调视觉语言模型,实现分层思维链推理
  • 在AI Habitat上比顶尖方法成功率高22.4%,路径加权成功率也更优
  • 适合研究视觉导航、具身智能和通用决策的学者与工程师

视觉目标导航(ObjectNav)要求机器人在未见过的环境中,基于第一人称观测定位目标物体。传统端到端学习方法因依赖记忆空间模式而非结构化推理,导致泛化能力差。本文提出闭合环路分层思维链导航(CL-CoTNav),一种基于视觉语言模型(VLM)的导航框架,融合结构化推理与闭环反馈。通过人类示范轨迹生成的多轮问答数据对VLM进行微调,支持分层思维链(H-CoT)提示,系统性提取组合知识以优化感知与决策,模仿人类逐步推理过程。进一步提出闭合环路H-CoT机制,将检测与推理置信度融入训练,自适应加权高置信度数据对,缓解噪声输入影响,增强对幻觉或错误推理的鲁棒性。在AI Habitat环境中的大量实验表明,该方法在未见场景和新物体类别上均显著优于现有最优方法,导航成功率(SR)与路径加权成功率(SPL)平均提升22.4%。相关数据集、模型及补充视频已公开。

原文摘要 · Abstract (English)

Visual Object Goal Navigation (ObjectNav) requires a robot to locate a target object in an unseen environment using egocentric observations. However, decision-making policies often struggle to transfer to unseen environments and novel target objects, which is the core generalization problem. Traditional end-to-end learning methods exacerbate this issue, as they rely on memorizing spatial patterns rather than employing structured reasoning, limiting their ability to generalize effectively. In this letter, we introduce Closed-Loop Hierarchical Chain-of-Thought Navigation (CL-CoTNav), a vision-language model (VLM)-driven ObjectNav framework that integrates structured reasoning and closed-loop feedback into navigation decision-making. To enhance generalization, we fine-tune a VLM using multi-turn question-answering (QA) data derived from human demonstration trajectories. This structured dataset enables hierarchical Chain-of-Thought (H-CoT) prompting, systematically extracting compositional knowledge to refine perception and decision-making, inspired by the human cognitive process of locating a target object through iterative reasoning steps. Additionally, we propose a Closed-Loop H-CoT mechanism that incorporates detection and reasoning confidence scores into training. This adaptive weighting strategy guides the model to prioritize high-confidence data pairs, mitigating the impact of noisy inputs and enhancing robustness against hallucinated or incorrect reasoning. Extensive experiments in the AI Habitat environment demonstrate CL-CoTNav's superior generalization to unseen scenes and novel object categories. Our method consistently outperforms state-of-the-art approaches in navigation success rate (SR) and success weighted by path length (SPL) by 22.4\%. We release our datasets, models, and supplementary videos on our project page.

视觉导航思维链具身智能VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。