arXiv:2605.19690cs.RO2026-05中稿 · the 2026 IEEE Inte…

让导航模型在新环境里学得快还不会忘老本领

D-CLING: Prior-Preserving Depth-Conditioned Fine-Tuning for Navigation Foundation Models

论文配图:D-CLING: Prior-Preserving Depth-Conditioned Fine-Tuning for Navigation Foundation Models
图 1 · 摘自论文原文
  • 用零初始化残差路径连接预训练主干,保留先验知识
  • 实测长距离导航碰撞少,人类干预降低40%以上
  • 适合需要持续学习的机器人导航场景

基于大规模跨体数据训练的导航基础模型(NFMs)展现出强大泛化能力。但在新环境或摄像头配置下进行领域内微调时,模型常出现避障能力差或无法抵达目标的问题。更严重的是,小样本微调会侵蚀预训练阶段积累的知识,削弱模型泛化性能。为此,本文提出D-CLING方法:借鉴ControlNet思想,通过零初始化残差路径附加可训练的预训练主干副本,使模型高效学习特定场景几何信息的同时,保持原有行为的先验知识。实验表明,该方法在真实世界导航任务中显著提升鲁棒性,实现低碰撞、低人工干预的长时程导航。离线分析进一步显示,该方法在微调数据集外仍能保持甚至提升动作预测能力,为通用导航的持续学习提供了关键洞见。

原文摘要 · Abstract (English)

Navigation Foundation Models (NFMs) trained on large cross-embodied datasets have demonstrated powerful generalizability in various scenarios. Adopting in-domain fine-tuning for an NFM efficiently calibrates the visuomotor policy, promising further improvement even in a novel scenario. However, the fine-tuned models still suffer from poor obstacle avoidance or fail to properly reach the provided goals. Furthermore, model updates using a small subset of data typically erode the pre-trained prior, compromising the pre-training generalization. Consequently, fine-tuning deteriorates the capability of the model for robust and accurate navigation. In this work, we present a novel fine-tuning method that leverages large-scale pre-training while efficiently learning in novel setups, such as environments or camera configurations. In particular, inspired by ControlNet, we fine-tune an NFM by attaching a trainable copy of the pre-trained backbone using zero-initialized residual pathways, thereby learning geometric cues. This design enables the model to efficiently acquire in-domain geometry while preserving pre-trained knowledge across various behaviors. Despite its simplicity, our comprehensive evaluation of real-world navigation suggests that our proposal effectively enables robust long-horizon navigation with minimal collisions and human intervention. Additionally, our offline analysis shows that the proposed method maintains or further improves action prediction capabilities beyond the fine-tuned dataset, providing a key insight into continual learning for general navigation. The project page: https://toyotafrc.github.io/DCLING-Proj/

导航模型微调持续学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。