让无人机在遮挡下也能精准追踪目标,靠空间感知与闭环控制。
CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking

- 引入空间感知机制,联合定位目标、判断可见性并生成连续飞行指令。
- 在未见测试集上,平均位移误差降低35.3%,追踪成功率提升29.8%。
- 适合需要复杂环境持久追踪的无人机应用,如城市安防与搜救。
动态目标追踪对在复杂城市环境中运行的无人机至关重要,因目标与摄像机视角持续变化。现有视觉-语言-动作(VLA)策略虽能有效追踪可见目标,但在建筑物、植被或路边物体遮挡时性能显著下降。长时间遮挡会导致策略丢失目标状态,执行错误动作,并通过后续观测放大误差,直至无法重新获取。为此,我们提出CosFly-VLA,一种具有空间感知能力的VLA模型,通过结构化预测接口联合定位目标、估计其可见性并生成连续飞行动作。训练采用大规模多源数据配方:在50万样本混合池上进行空间接地持续预训练(CPT),注入无人机视角深度、距离及三维空间推理能力;随后通过三阶段课程制监督微调(SFT),经多头热身与自然及困难/长时遮挡数据两阶段课程学习完成专业化;再以思维链(CoT)训练教授恢复导向推理路径,最终生成结构化答案;最后通过闭环强化学习(RL)优化追踪行为,奖励包含保持安全距离、定位质量、避障与任务成功。相比OpenVLA,CosFly-VLA-0.8B在已见测试集上开环平均位移误差(ADE)降低34.1%,未见测试集降低35.3%;闭环优化使成功率(SR)分别提升29.8%和2.5%。结果表明,该方法从可见帧模仿迈向基于空间感知的闭环动作控制,在共享真实状态历史下验证有效。
原文摘要 · Abstract (English)
Dynamic target tracking is essential for Unmanned Aerial Vehicles (UAVs) operating in complex urban environments, where both the target and the camera viewpoint change continuously. Existing Vision-Language-Action (VLA) policies can track visible targets effectively, but their performance often degrades when buildings, vegetation, or roadside objects block the line of sight. During sustained occlusion, a policy may lose the target state, execute actions toward an incorrect region, and amplify this error through subsequent observations until re-acquisition becomes impossible. To this end, we present CosFly-VLA, a spatially aware VLA model that jointly grounds the target, estimates its visibility, and generates continuous flight actions through a structured prediction interface. To train this policy, we use a large-scale recipe over diverse data sources. Spatially Grounded Continued Pretraining (CPT) on a 500k mixed pool injects UAV-view depth, distance, and 3-D spatial reasoning. A three-stage Curriculum-based Supervised Fine-Tuning (SFT) process then specializes the tracker through multi-head warm-up followed by two-stage curriculum learning over natural and hard / long-occlusion data. Chain-of-Thought (CoT) training subsequently teaches recovery-oriented reasoning traces before structured answers. Finally, a closed-loop Reinforcement Learning (RL) stage optimizes tracking behavior with a multi-component reward covering stand-off tracking, grounding quality, collision avoidance, and task success. Relative to OpenVLA, CosFly-VLA-0.8B reduces open-loop Average Displacement Error (ADE) by 34.1% on seen-test and 35.3% on unseen-test. Closed-loop optimization improves Success Rate (SR) by 29.8% and 2.5%, respectively. These results demonstrate progress from visible-frame imitation toward spatially grounded action-closed-loop control, evaluated under a shared oracle state history.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。