用视觉语言动作模型让无人机像人一样竞速,速度超2米/秒。
RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour
- 基于VLA模型,通过视觉语言指令实时调整飞行策略。
- 最高时速达2.02米/秒,平均速度1.04米/秒,高速下仍稳定操控。
- 相比OpenVLA和RT-2在运动与语义泛化上全面领先,适合竞速场景。
RaceVLA提出一种基于视觉-语言-动作(VLA)模型的自主竞速无人机导航方法,旨在模仿人类飞行员的行为。该模型在自建的竞速无人机数据集上微调,具备强泛化能力,即使在复杂环境中也能适应实时环境反馈。实验显示,其运动泛化得分75.0,显著高于OpenVLA的60.0;语义泛化45.5,优于36.3。尽管动态环境下视觉与物理泛化略有下降(分别为79.6和50.0),但仍全面超越RT-2(视觉79.6 vs 52.0,运动75.0 vs 55.0,物理50.0 vs 26.7,语义45.5 vs 38.8)。平均速度达1.04米/秒,最高速度2.02米/秒,证明其在高速竞速中表现稳健。代码、预训练权重与数据集已公开于https://racevla.github.io/。
原文摘要 · Abstract (English)
RaceVLA presents an innovative approach for autonomous racing drone navigation by leveraging Visual-Language-Action (VLA) to emulate human-like behavior. This research explores the integration of advanced algorithms that enable drones to adapt their navigation strategies based on real-time environmental feedback, mimicking the decision-making processes of human pilots. The model, fine-tuned on a collected racing drone dataset, demonstrates strong generalization despite the complexity of drone racing environments. RaceVLA outperforms OpenVLA in motion (75.0 vs 60.0) and semantic generalization (45.5 vs 36.3), benefiting from the dynamic camera and simplified motion tasks. However, visual (79.6 vs 87.0) and physical (50.0 vs 76.7) generalization were slightly reduced due to the challenges of maneuvering in dynamic environments with varying object sizes. RaceVLA also outperforms RT-2 across all axes - visual (79.6 vs 52.0), motion (75.0 vs 55.0), physical (50.0 vs 26.7), and semantic (45.5 vs 38.8), demonstrating its robustness for real-time adjustments in complex environments. Experiments revealed an average velocity of 1.04 m/s, with a maximum speed of 2.02 m/s, and consistent maneuverability, demonstrating RaceVLA's ability to handle high-speed scenarios effectively. These findings highlight the potential of RaceVLA for high-performance navigation in competitive racing contexts. The RaceVLA codebase, pretrained weights, and dataset are available at this http URL: https://racevla.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。