arXiv:2503.02723cs.RO2025-03中稿 · IROS 2025被引 10

用视觉语言模型让微型无人机群动态调整避障策略,应对复杂环境。

ImpedanceGPT: VLM-driven Impedance Control of Swarm of Mini-drones for Intelligent Navigation in Dynamic Environment

  • 结合视觉语言模型与检索增强生成,实现环境语义理解。
  • 对静态障碍物可速达1.4米/秒,遇人时降至0.7米/秒以保安全。
  • 适合需要智能协同避障的无人机集群场景,如救援或巡检。

群体机器人在动态不可预测环境中实现自主运行至关重要。然而,如何在充满动态生物(如人类)和动态非生物(如移动物体)障碍物的环境中确保安全高效导航仍是主要挑战。本文提出ImpedanceGPT,一种将视觉语言模型(VLM)与检索增强生成(RAG)结合的新系统,用于实现微型无人机群在复杂环境中的实时自适应导航。其核心创新在于融合VLM与RAG,使无人机具备更强的环境语义理解能力,能够根据障碍物类型和环境条件动态调节阻抗控制参数。该方法不仅保障了安全精准导航,还提升了群内协同效率。实验表明,该系统有效:在理想光照下,障碍物检测与检索准确率达80%;在静态环境中,无人机以1.4米/秒速度通过非生物动态障碍,遇人时则减速至0.7米/秒并增加间距;在动态环境中,靠近硬障碍物时速度降至1.0米/秒,面对移动人类则进一步降至0.6米/秒并增大偏移以确保安全避让。

原文摘要 · Abstract (English)

Swarm robotics plays a crucial role in enabling autonomous operations in dynamic and unpredictable environments. However, a major challenge remains ensuring safe and efficient navigation in environments filled with both dynamic alive (e.g., humans) and dynamic inanimate (e.g., non-living objects) obstacles. In this paper, we propose ImpedanceGPT, a novel system that combines a Vision-Language Model (VLM) with retrieval-augmented generation (RAG) to enable real-time reasoning for adaptive navigation of mini-drone swarms in complex environments. The key innovation of ImpedanceGPT lies in the integration of VLM and RAG, which provides the drones with enhanced semantic understanding of their surroundings. This enables the system to dynamically adjust impedance control parameters in response to obstacle types and environmental conditions. Our approach not only ensures safe and precise navigation but also improves coordination between drones in the swarm. Experimental evaluations demonstrate the effectiveness of the system. The VLM-RAG framework achieved an obstacle detection and retrieval accuracy of 80 % under optimal lighting. In static environments, drones navigated dynamic inanimate obstacles at 1.4 m/s but slowed to 0.7 m/s with increased separation around humans. In dynamic environments, speed adjusted to 1.0 m/s near hard obstacles, while reducing to 0.6 m/s with higher deflection to safely avoid moving humans.

无人机群视觉语言模型智能导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。