无需训练的无人机系统,能听懂指令并快速避障飞行
QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight
- 用异步多智能体架构分离决策与控制
- 实测可在室内以5米/秒速度飞行
- 适合需要实时响应的无人机导航场景
我们提出QuadAgent,一种无需训练的视觉语言引导四旋翼敏捷飞行代理系统。不同于以往端到端或串行代理方法,QuadAgent采用异步多智能体架构,将高层推理与底层控制解耦:前景工作代理处理当前任务和用户指令,背景代理执行前瞻推理。系统通过印象图(Impression Graph)维持场景记忆,该轻量级拓扑地图基于稀疏关键帧构建,并借助基于视觉的障碍物规避网络保障飞行安全。仿真结果表明,QuadAgent在效率和响应速度上优于基线方法。真实世界实验显示,它能理解复杂指令,推理周围环境,并在密集室内空间以最高5米/秒的速度导航。
原文摘要 · Abstract (English)
We present QuadAgent, a training-free agent system for agile quadrotor flight guided by vision-language inputs. Unlike prior end-to-end or serial agent approaches, QuadAgent decouples high-level reasoning from low-level control using an asynchronous multi-agent architecture: Foreground Workflow Agents handle active tasks and user commands, while Background Agents perform look-ahead reasoning. The system maintains scene memory via the Impression Graph, a lightweight topological map built from sparse keyframes, and ensures safe flight with a vision-based obstacle avoidance network. Simulation results show that QuadAgent outperforms baseline methods in efficiency and responsiveness. Real-world experiments demonstrate that it can interpret complex instructions, reason about its surroundings, and navigate cluttered indoor spaces at speeds up to 5 m/s.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。