让无人机仅靠机载感知实现自然语言导航,零样本迁移效果佳。
SINGER: An Onboard Generalist Vision-Language Navigation Policy for Drones
- 用高保真仿真生成数据,结合路径规划生成无碰撞导航示范
- 训练出轻量级端到端视觉运动策略,实时控制性能强
- 适合需要自主导航的无人机应用,尤其在未知环境表现优
大型视觉-语言模型推动了开放词汇机器人策略的发展,如通用机器人操作策略,使机器人能完成自然语言描述的复杂任务。然而,由于缺乏大规模演示数据、无人机对实时控制的高要求以及可靠外部定位模块的缺失,开放世界中的语言引导无人机自主导航仍是未解难题。本文提出SINGER,一种仅依赖机载感知与计算的语言引导无人机自主导航通用策略。为训练鲁棒的开放词汇导航策略,SINGER采用三个核心组件:(i) 基于高斯点云的逼真语言嵌入飞行仿真器,实现高效数据生成且模拟到现实的差距极小;(ii) 受RRT启发的多轨迹生成专家,用于生成无碰撞导航示范;(iii) 轻量级端到端视觉运动策略,支持实时闭环控制。通过大量硬件飞行实验,我们验证了该策略在未见过环境和未见过语言目标下的卓越零样本模拟到现实迁移能力。在约70万至100万条观测-动作对上训练并部署于真实硬件时,相比速度控制的语义引导基线,SINGER平均成功率提升23.33%,目标保持在视场内时间增加16.67%,碰撞次数减少10%。
原文摘要 · Abstract (English)
Large vision-language models have driven remarkable progress in open-vocabulary robot policies, e.g., generalist robot manipulation policies, that enable robots to complete complex tasks specified in natural language. Despite these successes, open-vocabulary autonomous drone navigation remains an unsolved challenge due to the scarcity of large-scale demonstrations, real-time control demands of drones for stabilization, and lack of reliable external pose estimation modules. In this work, we present SINGER for language-guided autonomous drone navigation in the open world using only onboard sensing and compute. To train robust, open-vocabulary navigation policies, SINGER leverages three central components: (i) a photorealistic language-embedded flight simulator with minimal sim-to-real gap using Gaussian Splatting for efficient data generation, (ii) an RRT-inspired multi-trajectory generation expert for collision-free navigation demonstrations, and these are used to train (iii) a lightweight end-to-end visuomotor policy for real-time closed-loop control. Through extensive hardware flight experiments, we demonstrate superior zero-shot sim-to-real transfer of our policy to unseen environments and unseen language-conditioned goal objects. When trained on ~700k-1M observation action pairs of language conditioned visuomotor data and deployed on hardware, SINGER outperforms a velocity-controlled semantic guidance baseline by reaching the query 23.33% more on average, and maintains the query in the field of view 16.67% more on average, with 10% fewer collisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。