arXiv:2409.13337cs.RO2024-09被引 2

让无人机在目标不可见时仍能通过生成模型导航到指定视角

Invisible Servoing: a Visual Servoing Approach with Return-Conditioned Latent Diffusion

  • 用跨模态VAE提取图像隐表示,结合扩散模型生成导航轨迹
  • 在初始视图无目标时仍可成功到达目标视角,成功率100%
  • 适合做视觉伺服导航的无人机系统,尤其适用于遮挡场景

本文提出一种基于潜在去噪扩散概率模型(DDPM)的新型视觉伺服(VS)方法,探索生成模型在无人机(UAV)视觉导航中的应用。与传统方法不同,该方法可在目标初始不可见时仍引导无人机到达期望目标视角。其关键在于学习一个用于规划的潜在表示,并利用包含目标不可见初始视图的轨迹数据集进行训练。通过跨模态变分自编码器从原始图像中提取紧凑的隐表示,当前图像输入后,DDPM在潜空间生成驱动机器人平台到达目标视觉位置的轨迹。该方法在两种通用多旋翼无人机(四旋翼和六旋翼)的仿真环境中验证,结果表明即使初始视图中目标不可见,也能成功抵达目标视角。

原文摘要 · Abstract (English)

In this paper, we present a novel visual servoing (VS) approach based on latent Denoising Diffusion Probabilistic Models (DDPMs), that explores the application of generative models for vision-based navigation of UAVs (Uncrewed Aerial Vehicles). Opposite to classical VS methods, the proposed approach allows reaching the desired target view, even when the target is initially not visible. This is possible thanks to the learning of a latent representation that the DDPM uses for planning and a dataset of trajectories encompassing target-invisible initial views. A compact representation is learned from raw images using a Cross-Modal Variational Autoencoder. Given the current image, the DDPM generates trajectories in the latent space driving the robotic platform to the desired visual target. The approach has been validated in simulation using two generic multi-rotor UAVs (a quadrotor and a hexarotor). The results show that we can successfully reach the visual target, even if not visible in the initial view.

视觉伺服扩散模型无人机导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。