arXiv:2510.16624cs.CVcs.RO2025-10

用视觉实现低成本无人机自主飞行,精准测距避障

Self-Supervised Learning to Fly using Efficient Semantic Segmentation and Metric Depth Estimation for Low-Cost Autonomous UAVs

  • 融合语义分割与单目深度估计,无须GPS或激光雷达
  • 自适应尺度算法使非度量深度误差仅14.4厘米
  • 轻量化模型实现实时分割,适合资源受限无人机

本文提出一种仅依赖视觉的自主飞行系统,适用于小型无人机在受控室内环境运行。系统结合语义分割与单目深度估计,实现障碍物避让、场景探索及安全着陆,无需GPS或昂贵传感器(如LiDAR)。核心创新在于自适应尺度因子算法,利用语义地面检测和相机内参将非度量深度预测转换为精确度量距离,均方距离误差达14.4厘米。采用知识蒸馏框架,由基于颜色的支持向量机(SVM)教师生成训练数据,供轻量级U-Net学生网络(160万参数)进行实时语义分割;复杂环境可替换为先进分割模型。测试在5×4米实验室环境中进行,设8个纸板障碍物模拟城市结构。30次真实飞行测试与100次数字孪生环境测试表明,该方法显著提升巡检航程、缩短任务时间,且成功率保持100%。通过端到端学习,紧凑的学生网络从最优方法生成的演示数据中学习完整飞行策略,实现87.5%的自主任务成功率。本工作推进了结构化环境中实用的视觉导航技术,解决了度量深度估计与计算效率难题,助力资源受限平台部署。

原文摘要 · Abstract (English)

This paper presents a vision-only autonomous flight system for small UAVs operating in controlled indoor environments. The system combines semantic segmentation with monocular depth estimation to enable obstacle avoidance, scene exploration, and autonomous safe landing operations without requiring GPS or expensive sensors such as LiDAR. A key innovation is an adaptive scale factor algorithm that converts non-metric monocular depth predictions into accurate metric distance measurements by leveraging semantic ground plane detection and camera intrinsic parameters, achieving a mean distance error of 14.4 cm. The approach uses a knowledge distillation framework where a color-based Support Vector Machine (SVM) teacher generates training data for a lightweight U-Net student network (1.6M parameters) capable of real-time semantic segmentation. For more complex environments, the SVM teacher can be replaced with a state-of-the-art segmentation model. Testing was conducted in a controlled 5x4 meter laboratory environment with eight cardboard obstacles simulating urban structures. Extensive validation across 30 flight tests in a real-world environment and 100 flight tests in a digital-twin environment demonstrates that the combined segmentation and depth approach increases the distance traveled during surveillance and reduces mission time while maintaining 100% success rates. The system is further optimized through end-to-end learning, where a compact student neural network learns complete flight policies from demonstration data generated by our best-performing method, achieving an 87.5% autonomous mission success rate. This work advances practical vision-based drone navigation in structured environments, demonstrating solutions for metric depth estimation and computational efficiency challenges that enable deployment on resource-constrained platforms.

视觉导航无人机深度估计轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。