arXiv:2509.02983cs.ROcs.CV2025-09被引 1

用扩散模型实现无地图水下视觉导航,靠迁移深度特征提升泛化能力。

DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features

  • 通过两阶段训练:先在空中数据上预训练,再用水下深度任务微调特征提取器。
  • 在模拟与真实水下环境均实现安全避障和地形感知,4自由度运动控制稳定。
  • 适合水下机器人、无人艇等需要自主导航的场景,尤其适合缺乏地图的未知水域。

自主水下导航因传感能力受限及难以构建精准地图而极具挑战。本文提出基于扩散模型的水下视觉导航策略DUViN,实现无需预建地图的视觉端到端4自由度运动控制。DUViN可引导潜水器避开障碍物,并保持与地形的安全感知高度。为应对大规模水下导航数据难获取的问题,我们利用深度特征并引入新颖的模型迁移策略,确保从空中到水下的域转移中具备鲁棒泛化能力。训练分两阶段:首先在空中数据集上使用预训练深度特征提取器训练扩散导航策略;其次在水下深度估计任务上重训练提取器,并将其集成到第一阶段训练好的导航策略中。仿真与真实水下环境实验验证了方法的有效性与泛化能力。视频演示见:https://www.youtube.com/playlist?list=PLqt2s-RyCf1gfXJgFzKjmwIqYhrP4I-7Y。

原文摘要 · Abstract (English)

Autonomous underwater navigation remains a challenging problem due to limited sensing capabilities and the difficulty of constructing accurate maps in underwater environments. In this paper, we propose a Diffusion-based Underwater Visual Navigation policy via knowledge-transferred depth features, named DUViN, which enables vision-based end-to-end 4-DoF motion control for underwater vehicles in unknown environments. DUViN guides the vehicle to avoid obstacles and maintain a safe and perception awareness altitude relative to the terrain without relying on pre-built maps. To address the difficulty of collecting large-scale underwater navigation datasets, we propose a method that ensures robust generalization under domain shifts from in-air to underwater environments by leveraging depth features and introducing a novel model transfer strategy. Specifically, our training framework consists of two phases: we first train the diffusion-based visual navigation policy on in-air datasets using a pre-trained depth feature extractor. Secondly, we retrain the extractor on an underwater depth estimation task and integrate the adapted extractor into the trained navigation policy from the first step. Experiments in both simulated and real-world underwater environments demonstrate the effectiveness and generalization of our approach. The experimental videos are available at https://www.youtube.com/playlist?list=PLqt2s-RyCf1gfXJgFzKjmwIqYhrP4I-7Y.

水下导航扩散模型视觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。