arXiv:2504.21695cs.ROcs.AI2025-04被引 2

无需外部数据,用单目摄像头和飞控数据自监督训练无人机模型,提升高速近障飞行精度。

Self-Supervised Monocular Visual Drone Model Identification through Improved Occlusion Handling

  • 自监督学习结合单目视频与飞控数据,训练神经网络无人机模型。
  • 新遮挡处理方法使里程计误差平均降低15%,高精度支持高速飞行。
  • 可部署于任意无人机,适用于真实复杂环境,无需昂贵标注数据。

在无GPS环境下,自运动估计对无人机飞行至关重要。视觉方法在高速飞行或近距离障碍物下易受运动模糊和大范围遮挡影响。现有方案通常依赖外部运动捕捉系统提供的真值数据,以监督方式学习无人机模型,难以扩展至不同环境和机型。本文提出一种自监督学习框架,仅使用机载单目视频、惯性测量单元(IMU)和电机反馈数据,训练基于神经网络的无人机模型。首先训练一个自监督相对位姿估计模型作为教师,再通过其指导学生无人机模型的学习。为应对高速近障场景,我们改进了遮挡处理机制,使最终里程计的均方根误差平均降低15%。所获学生模型能从机载数据中成功提取,且在更高飞行速度下比教师模型更精确。将该神经无人机模型集成到传统滤波式视觉惯性里程计系统(ROVIO)中,在靠近障碍物的激进3D竞速轨迹上显著提升了定位精度。自监督自运动估计为实现从受控实验室环境向真实世界无人机应用的跨越迈出关键一步,视觉与无人机模型融合将推动任意无人机在任何环境中实现更高时速飞行与更优状态估计。

原文摘要 · Abstract (English)

Ego-motion estimation is vital for drones when flying in GPS-denied environments. Vision-based methods struggle when flight speed increases and close-by objects lead to difficult visual conditions with considerable motion blur and large occlusions. To tackle this, vision is typically complemented by state estimation filters that combine a drone model with inertial measurements. However, these drone models are currently learned in a supervised manner with ground-truth data from external motion capture systems, limiting scalability to different environments and drones. In this work, we propose a self-supervised learning scheme to train a neural-network-based drone model using only onboard monocular video and flight controller data (IMU and motor feedback). We achieve this by first training a self-supervised relative pose estimation model, which then serves as a teacher for the drone model. To allow this to work at high speed close to obstacles, we propose an improved occlusion handling method for training self-supervised pose estimation models. Due to this method, the root mean squared error of resulting odometry estimates is reduced by an average of 15%. Moreover, the student neural drone model can be successfully obtained from the onboard data. It even becomes more accurate at higher speeds compared to its teacher, the self-supervised vision-based model. We demonstrate the value of the neural drone model by integrating it into a traditional filter-based VIO system (ROVIO), resulting in superior odometry accuracy on aggressive 3D racing trajectories near obstacles. Self-supervised learning of ego-motion estimation represents a significant step toward bridging the gap between flying in controlled, expensive lab environments and real-world drone applications. The fusion of vision and drone models will enable higher-speed flight and improve state estimation, on any drone in any environment.

无人机自监督学习视觉里程计姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。