用扩散模型解决像素目标导航中的3D定位与避障难题
OccPlanner: Goal-Aware Occupancy-Conditioned Diffusion Planner for Pixel-Goal Navigation

- 基于视觉上下文和3D占用特征逐步优化目标表示
- 在5-8米距离下成功率从20.81%提升至71.55%
- 适合需要精准像素级导航的机器人场景
像素目标导航直接以代理相机视图指定目标,但目标像素不提供度量深度或可通行性,导致3D目标定位与无碰撞连续规划极具挑战。我们提出OccPlanner,一种目标感知的占用条件扩散规划器,将像素目标在自身坐标系中进行度量空间定位,并依次结合时序视觉上下文和学习到的局部3D占用特征来条件化目标表示。为大规模提供占用监督,我们引入L3ROcc,通过几何重建与基于射线的可见性推理,将单目RGB导航视频转换为机器人中心的局部3D占用标注。我们在InternData-N1上训练OccPlanner,并在闭环仿真中评估其在InternScenes四个未见场景类别及两个目标距离范围的表现。在5-8米设置下,相较于NavDP,OccPlanner在四类场景上的平均成功率达71.55%,在杂乱-易和杂乱-难场景中分别达到86.20%和84.92%。在Unitree Go2上的真实世界开环实验进一步验证了基于L3ROcc生成监督的模拟到现实迁移与适应能力。
原文摘要 · Abstract (English)
Pixel-goal navigation specifies targets directly in the agent's camera view, but a target pixel provides neither metric depth nor traversability, making 3D goal grounding and collision-free continuous planning challenging. We present OccPlanner, a goal-aware occupancy-conditioned diffusion planner that grounds pixel goals in egocentric metric space and sequentially conditions the goal representation on temporal visual context and learned local 3D occupancy features. To provide occupancy supervision at scale, we introduce L3ROcc, which converts monocular RGB navigation videos into robot-centric local 3D occupancy annotations through geometric reconstruction and ray-based visibility reasoning. We train OccPlanner on InternData-N1 and evaluate it in closed-loop simulation across four unseen scene categories from InternScenes and two goal-distance ranges. In the 5-8 m setting, OccPlanner increases the average success rate (SR) over NavDP from 20.81% to 71.55% across the four categories, reaching 86.20% and 84.92% in cluttered-easy and cluttered-hard scenes, respectively. Real-world open-loop experiments on a Unitree Go2 further provide initial evidence of sim-to-real transfer and adaptation with L3ROcc-generated supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。