用轻量模型+自动优化,让无人机在复杂环境导航更安全高效
AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

- 用20亿参数轻量模型结合高精度视觉输入,替代70亿参数大模型
- 在未测绘场景中成功率达49.16%,碰撞率显著降低
- 无需人工标注,通过物理仿真回滚自动生成高质量训练数据
无人机视觉语言导航需在复杂三维环境中实现快速响应控制。现有极简端到端范式虽具潜力,但普遍依赖含数十亿参数的大型语言模型,导致实时边缘部署延迟过高。本文挑战这一参数密集依赖。跨尺度评估揭示关键洞见:感知质量远超语言推理能力。我们证明,配备高保真视觉输入的20亿参数模型可完全匹配70亿基线模型的整体成功率。然而,该极简策略暴露纯行为克隆(BC)固有的鲁棒性缺陷——缺乏显式负反馈,导致代理无法内化稳健的空间约束,在分布外(OOD)场景中碰撞率惊人。为克服此脆弱性而不依赖不可扩展的人工标注,我们提出AeroDPO,一种零成本自动化直接偏好优化流程,由确定性物理仿真状态回滚驱动。检测到碰撞后,系统自主回溯环境以提取因果推理错误作为被拒绝动作,采用解耦特权干预合成避障优选动作,并借助离线视觉语言检查器过滤视觉模糊。通过为20亿模型注入此自动化数据飞轮,AeroDPO在未测绘场景中将成功率提升至49.16%,同时大幅抑制碰撞率,确立自主空中代理新SOTA。
原文摘要 · Abstract (English)
Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billions of parameters, incurring prohibitive latency for real-world edge deployment. In this paper, we challenge this parameter-heavy reliance. Comprehensive cross-scale evaluations reveal the critical insight that perception quality fundamentally outweighs language reasoning capacity. We demonstrate that a lightweight 2B model equipped with high-fidelity visual inputs completely matches the overall success rates of massive 7B baselines. However, this minimalist policy exposes a fundamental robustness flaw inherent to pure Behavior Cloning (BC). Lacking explicit negative feedback, the agent fails to internalize robust spatial constraints and exhibits alarming collision rates in out-of-distribution (OOD) scenarios. To overcome this vulnerability without relying on unscalable human annotations, we propose AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback. Upon detecting collisions, the system autonomously rewinds the environment to extract causal reasoning errors as rejected actions, applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers, and leverages an offline vision language inspector to filter visual ambiguities. By equipping our 2B model with this automated data flywheel, AeroDPO boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。