arXiv:2606.05506cs.CV2026-06中稿 · ICRA

用激光雷达引导视觉表征学习,提升导航模型跨场景泛化能力

Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning

论文配图:Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning
图 1 · 摘自论文原文
  • 用激光雷达指导对比学习,让视觉特征聚焦导航结构而非外观细节
  • 在多种室内外环境中,部署时仅用单目图像仍能实现优异跨场景迁移
  • 适合需要强泛化能力的机器人导航研究,尤其在光照/语义变化大的场景

我们提出一种传感器引导的自适应对比学习框架,用于点目标导航中的视觉表征学习。训练阶段,特权激光雷达通过几何感知相似性度量和自适应温度缩放,指导对比目标,促使视觉嵌入捕捉与导航相关的结构信息,而非场景特定的外观特征。所得到的编码器可独立预训练、冻结后作为强化学习的感知主干,实现表征学习与策略优化解耦。我们进一步引入表征预训练与策略学习间的跨阶段域差异,抑制环境特异性捷径,促进对任务相关特征的依赖。大量高保真仿真实验表明,该方法显著提升了策略层面的跨场景迁移性能。部署时,智能体仅依赖单目RGB观测及标准任务输入(如目标位置、本体感知信号),无需访问激光雷达或其他特权传感器。相比大型预训练视觉模型和标准对比基线,本方法在严重外观与语义变化下表现更优。我们还发布了多模态数据集,以支持未来关于特权引导视觉表征学习的研究。代码已开源。

原文摘要 · Abstract (English)

We propose a sensor-guided adaptive contrastive learning framework for visual representation learning in PointGoal navigation. During training, privileged LiDAR sensing guides the contrastive objective through a geometry-aware similarity metric and adaptive temperature scaling, encouraging visual embeddings to capture navigation-relevant structure rather than scene-specific appearance. The resulting encoder is pretrained independently, frozen, and used as the perceptual backbone for reinforcement learning, decoupling representation learning from policy optimization. We further introduce a cross-stage domain mismatch between representation pretraining and policy learning to suppress environment-specific shortcuts and promote reliance on task-relevant features. Extensive experiments in high-fidelity simulation demonstrate that our approach significantly improves policy-level scene transfer across diverse indoor and outdoor environments. At deployment, the agent relies only on monocular RGB observations together with standard task-related inputs such as goal position and proprioceptive signals, without access to LiDAR or other privileged sensors. Our method outperforms large pretrained vision models and standard contrastive baselines under severe appearance and semantic shifts. We also release a multimodal dataset to support future research on privileged-guided visual representation learning for navigation. The code is available at:

导航对比学习多模态鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。