CoDrone让无人机在资源受限下实现高效自主导航,结合云端与边缘计算。
CoDrone: Autonomous Drone Navigation Assisted by Edge and Cloud Foundation Models
- 采用灰度图像与云边协同架构,降低计算与传输开销。
- 飞行距离提升40%,导航质量提高5%,适应不同速度与网络条件。
- 引入视觉语言模型和强化学习调度,增强复杂场景下的智能决策能力。
无人飞行器自主导航面临机载计算资源有限的挑战,导致部署的深度神经网络结构浅显,难以应对复杂环境。将任务卸载至远程边缘服务器又引入高延迟,造成系统设计中的固有权衡。为此,我们提出CoDrone——首个将基础模型融入无人机巡航场景的云-边-端协同计算框架,有效利用基础模型提升资源受限平台性能。为减少机载计算与数据传输负担,CoDrone采用灰度图像进行导航建模。当需增强环境感知时,借助边缘辅助的基础模型Depth Anything V2进行深度估计,并引入基于一维占用栅格的新型导航方法,实现细粒度场景理解的同时提升效率与表示简洁性。CoDrone的核心组件是基于深度强化学习的神经调度器,可无缝融合深度估计与自主导航决策,实现实时动态环境适应。此外,框架引入专用于无人机的视觉语言交互模块,结合领域定制的低级飞行原语,促进云端基础模型与无人机的有效互动,提升复杂未见场景下的开放集推理能力。实验结果表明,CoDrone在不同飞行速度与网络条件下均优于基线方法,平均飞行距离提升40%,平均导航质量提高5%。
原文摘要 · Abstract (English)
Autonomous navigation for Unmanned Aerial Vehicles faces key challenges from limited onboard computational resources, which restrict deployed deep neural networks to shallow architectures incapable of handling complex environments. Offloading tasks to remote edge servers introduces high latency, creating an inherent trade-off in system design. To address these limitations, we propose CoDrone - the first cloud-edge-end collaborative computing framework integrating foundation models into autonomous UAV cruising scenarios - effectively leveraging foundation models to enhance performance of resource-constrained unmanned aerial vehicle platforms. To reduce onboard computation and data transmission overhead, CoDrone employs grayscale imagery for the navigation model. When enhanced environmental perception is required, CoDrone leverages the edge-assisted foundation model Depth Anything V2 for depth estimation and introduces a novel one-dimensional occupancy grid-based navigation method - enabling fine-grained scene understanding while advancing efficiency and representational simplicity of autonomous navigation. A key component of CoDrone is a Deep Reinforcement Learning-based neural scheduler that seamlessly integrates depth estimation with autonomous navigation decisions, enabling real-time adaptation to dynamic environments. Furthermore, the framework introduces a UAV-specific vision language interaction module incorporating domain-tailored low-level flight primitives to enable effective interaction between the cloud foundation model and the UAV. The introduction of VLM enhances open-set reasoning capabilities in complex unseen scenarios. Experimental results show CoDrone outperforms baseline methods under varying flight speeds and network conditions, achieving a 40% increase in average flight distance and a 5% improvement in average Quality of Navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。