轻量级无人机可听懂指令自主飞行,无需外部设备支持。
GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics
- 用3D高斯点云模拟器+可微强化学习训练视觉语言动作模型。
- 仿真中未见任务成功率75%,真实硬件上达50%。
- 专家混合架构提升泛化能力,适合边缘部署的智能飞行系统。
能够在非结构化环境中理解并执行高层自然语言指令的自主无人机仍是长期目标。现有方法受限于手工设计技能、大量参数调优或计算密集型模型,难以在机载设备上运行。我们提出GRaD-Nav++,一种全机载运行、实时响应自然语言指令的轻量级视觉-语言-动作(VLA)框架。策略在基于3D高斯溅射(3DGS)的逼真三维模拟器中,通过可微强化学习(DiffRL)训练,实现从视觉和语言输入中高效学习底层控制。核心采用专家混合(MoE)动作头,动态路由计算以增强泛化性并缓解遗忘问题。多任务泛化实验显示,仿真中已训练任务成功率达83%,未见任务为75%;真实硬件部署后,已训练任务成功率为67%,未见任务为50%。多环境适应实验中,仿真平均成功率为81%,真实场景为67%。这些结果确立了全机载VLA飞行的新基准,表明紧凑高效的模型可在无外部基础设施依赖下实现可靠的语言引导导航。
原文摘要 · Abstract (English)
Autonomous drones capable of interpreting and executing high-level language instructions in unstructured environments remain a long-standing goal. Yet existing approaches are constrained by their dependence on hand-crafted skills, extensive parameter tuning, or computationally intensive models unsuitable for onboard use. We introduce GRaD-Nav++, a lightweight Vision-Language-Action (VLA) framework that runs fully onboard and follows natural-language commands in real time. Our policy is trained in a photorealistic 3D Gaussian Splatting (3DGS) simulator via Differentiable Reinforcement Learning (DiffRL), enabling efficient learning of low-level control from visual and linguistic inputs. At its core is a Mixture-of-Experts (MoE) action head, which adaptively routes computation to improve generalization while mitigating forgetting. In multi-task generalization experiments, GRaD-Nav++ achieves a success rate of 83% on trained tasks and 75% on unseen tasks in simulation. When deployed on real hardware, it attains 67% success on trained tasks and 50% on unseen ones. In multi-environment adaptation experiments, GRaD-Nav++ achieves an average success rate of 81% across diverse simulated environments and 67% across varied real-world settings. These results establish a new benchmark for fully onboard Vision-Language-Action (VLA) flight and demonstrate that compact, efficient models can enable reliable, language-guided navigation without relying on external infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。