arXiv:2606.08029cs.RO2026-06

从人类示范中学习类人目标导航,提升机器人探索效率。

IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations

论文配图:IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations
图 1 · 摘自论文原文
  • 基于人类示范提取搜索意图,构建空间-视觉候选集。
  • 在三个基准上达最优,零样本适配轮式、四足和人形机器人。
  • 无需微调即可跨平台迁移,适合多机器人系统部署。

目标导航要求机器人在未知环境中,通过部分观测决定下一步探索位置。高效搜索需模仿人类行为:有选择地探测视觉前景区域,并利用空间记忆避免重复访问。我们提出IntentNav,一种从人类示范中学习类人导航策略的空间-视觉模仿框架。为从低级动作推断高层搜索意图,引入基于前景的人类意图标注方法,通过前瞻分析人类示范,标记最能解释其未来搜索方向的前景区域。构建空间-视觉候选空间,其中鸟瞰图(BEV)记忆追踪已探索区域、未探索前景及轨迹历史,而自身视角视觉记忆为每个候选提供语义线索。使用视觉语言模型(VLM)策略从这些具身候选中选择,通过意图对齐目标鼓励一致且类人的探索行为。IntentNav在MP3D、HM3D-v1和HM3D-v2目标导航基准上达到当前最佳性能。所提出的候选级导航接口可零样本迁移至轮式、四足和人形机器人,无需进一步微调。

原文摘要 · Abstract (English)

Object navigation requires a robot to search for an unobserved target in an unknown environment by deciding where to explore next under partial observability. Effective search resembles human-like exploration: selectively probing visually promising frontiers while relying on spatial memory to avoid redundant revisits. We propose IntentNav, a spatial-visual imitation framework that learns human-like ObjectNav policies from human demonstrations. To infer high-level search intent from low-level human actions, we introduce Frontier-based Human-Intent Labeling, which looks ahead in human demonstrations and labels the frontier that best explains the demonstrator's future search direction. We construct a spatial-visual candidate space, where BEV memory tracks explored regions, unexplored frontiers, and trajectory history, while egocentric visual memory provides semantic cues for each candidate. A VLM policy is trained to select among these grounded candidates, using Intent-Aligned Objective to encourage consistent and human-like exploration. IntentNav achieves state-of-the-art performance on the MP3D, HM3D-v1 and HM3D-v2 ObjectNav benchmarks. The proposed candidate-level navigation interface transfers zero-shot to wheeled, quadruped, and humanoid robots without further VLM fine-tuning. \href{https://anonymous.4open.science/w/IntentNav/}{Project page}.

目标导航模仿学习多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。