arXiv:2608.30242cs.RO2026-08

让机器人在不同摄像头下也能稳定导航,摆脱镜头差异干扰。

CanonNav: Disentangling Navigation Behavior from Camera Geometry in Cross-Platform Visual Navigation

论文配图:CanonNav: Disentangling Navigation Behavior from Camera Geometry in Cross-Platform Visual Navigation
图 1 · 摘自论文原文
  • 通过归一化摄像头几何,分离导航行为与视角差异。
  • 结合伪标签监督,提升安全性和前进方向的准确性。
  • 跨平台训练,仅用RGB图像仍超越传统深度感知方法。

尽管视觉导航已通过跨平台示范数据的模仿学习取得进展,但充分挖掘此类数据仍具挑战。首先,直接从图像-轨迹对中学习会将导航行为与平台相关的摄像头几何纠缠在一起,导致策略需隐式推断摄像头参数,这是本质上难以求解的问题。其次,模仿学习虽能捕捉专家选择的动作序列,但无法揭示动作背后的中间决策过程。为此,我们提出CanonNav,一种将导航行为与摄像头几何解耦,并引入互补规划监督的视觉导航框架。CanonNav引入摄像头几何归一化,将视觉观测和轨迹转换到统一的相机一致表示空间。基于该表示,我们利用离线可通行性估计器生成伪标签,构建安全性和局部进展监督:前者惩罚不安全路径,后者指导机器人应朝何处前进。在多种摄像头配置和环境下的实验表明,即使推理时仅使用RGB图像,CanonNav在复杂场景中仍持续优于基于RGB的基线方法,甚至超过基于RGB-D的方法。

原文摘要 · Abstract (English)

While visual navigation has advanced through imitation learning from cross-platform demonstrations, fully leveraging such data remains challenging. First, directly learning from image-trajectory pairs entangles navigation behavior with platform-dependent camera geometry. This hinders consistent learning by forcing the policy to implicitly infer camera geometry from visual observations, an inherently ill-posed problem. Second, imitation learning from demonstrated trajectories captures the expert's chosen motion but leaves the intermediate decisions underlying that motion implicit. To address these issues, we propose CanonNav, a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations. CanonNav introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space. Building on this representation, we derive safety and local-progress supervision using pseudo-labels from an offline traversability estimator. Safety supervision penalizes unsafe trajectories, while local-progress supervision guides where the robot should advance. Experiments across diverse camera configurations and environments show that, despite using only RGB at inference, CanonNav consistently outperforms RGB-based baselines and even surpasses RGB-D-based methods in challenging scenarios.

视觉导航跨平台摄像头归一化模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。