轻量化视觉语言动作框架,让无人机在复杂环境自主飞行更安全高效
VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- 用3D高斯泼溅构建高保真数据集,解决真实与仿真差距
- 三阶段训练提升场景理解与长程导航能力,单任务成功率98.1%
- 实时轻量级动作模块+几何安全校正,适配资源受限无人机
本文提出VLA-AN,一种高效且可部署于机载端的视觉-语言-动作框架,专为复杂环境中无人机自主导航设计。针对现有大型导航模型存在的数据域差异、时序推理不足、生成式动作策略安全隐患及机载部署限制四大问题,提出四项改进:首先,采用3D高斯泼溅(3D-GS)构建高保真数据集,有效弥合域间差距;其次,设计渐进式三阶段训练框架,依次强化场景理解、核心飞行技能与复杂导航能力;第三,开发轻量级实时动作模块,结合几何安全校正,确保快速、无碰撞、稳定的指令生成,缓解随机生成策略的安全风险;最后,通过深度优化机载部署流程,使VLA-AN在资源受限无人机上实现推理吞吐量8.3倍提升。大量实验表明,VLA-AN显著提升空间定位精度、场景推理能力与长时程导航性能,单任务最高成功率达98.1%,为轻量化空中机器人实现全链路闭环自主提供高效实用方案。
原文摘要 · Abstract (English)
This paper proposes VLA-AN, an efficient and onboard Vision-Language-Action (VLA) framework dedicated to autonomous drone navigation in complex environments. VLA-AN addresses four major limitations of existing large aerial navigation models: the data domain gap, insufficient temporal navigation with reasoning, safety issues with generative action policies, and onboard deployment constraints. First, we construct a high-fidelity dataset utilizing 3D Gaussian Splatting (3D-GS) to effectively bridge the domain gap. Second, we introduce a progressive three-stage training framework that sequentially reinforces scene comprehension, core flight skills, and complex navigation capabilities. Third, we design a lightweight, real-time action module coupled with geometric safety correction. This module ensures fast, collision-free, and stable command generation, mitigating the safety risks inherent in stochastic generative policies. Finally, through deep optimization of the onboard deployment pipeline, VLA-AN achieves a robust real-time 8.3x improvement in inference throughput on resource-constrained UAVs. Extensive experiments demonstrate that VLA-AN significantly improves spatial grounding, scene reasoning, and long-horizon navigation, achieving a maximum single-task success rate of 98.1%, and providing an efficient, practical solution for realizing full-chain closed-loop autonomy in lightweight aerial robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。