arXiv:2605.17776cs.RO2026-05被引 1

构建首个城市无人机追踪大规模多模态数据集,解决动态目标跟踪数据缺失问题。

CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization

论文配图:CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization
图 1 · 摘自论文原文
  • 基于多约束优化生成2.4万步真实无人机轨迹,保证可视性与运动可行性。
  • 在7个视觉语言模型上提升跟踪准确率至78.3%~95.6%,超越零样本基线53~69个百分点。
  • 适合研究无人机自主导航、多模态感知与动态目标追踪的开发者使用。

近年来,空中视觉语言导航(VLN)数据集迅速增长,但主要聚焦于前往静态目的地的导航任务,而无人机视觉追踪——即持续跟随移动目标并保持可见性——仍缺乏专用训练数据。本文提出CosFlyTrack,一个面向城市环境的大型多模态无人机视觉追踪数据集及可扩展生成流程。该数据集包含约12,000条专家与扰动轨迹,源自6,000条行人路径,共240万时间步(约334小时),涵盖七种对齐数据通道:RGB图像、度量深度、语义分割、六自由度无人机位姿、含可见性标记的目标状态、中英文双语指令及轨迹对元数据。为生成高质量专家轨迹,我们设计了MuCO多约束优化器,直接在三维连续空间中规划,利用BVH加速碰撞与可视性查询,联合优化目标可见性、视角质量、避障、平滑性和运动学可行性,避免网格规划中的离散化误差与后处理平滑。在七个视觉语言模型上的微调实验表明,使用CosFlyTrack可将跟踪性能提升至SR@1米78.3%至95.6%,相比零样本基线提高53至69个百分点,验证其作为动态目标追踪训练资源的有效性。数据集已公开发布于https://huggingface.co/datasets/AutelRobotics/CosFly;评估脚本与预训练检查点托管于https://huggingface.co/AutelRobotics/CosFly-Track。

原文摘要 · Abstract (English)

Recent aerial vision-language navigation (VLN) datasets have grown rapidly, but they primarily address goal-oriented navigation to static destinations, leaving UAV visual tracking -- continuously following a moving target while maintaining visibility -- largely without dedicated training data. We introduce CosFlyTrack, a large-scale multi-modal dataset and scalable generation pipeline for UAV visual tracking in urban environments. The dataset provides approximately 12,000 expert and perturbed UAV trajectories generated from 6,000 pedestrian paths, comprising 2.4 million timesteps (approximately 334 hours) with seven aligned data channels: RGB, metric depth, semantic segmentation, six-degree-of-freedom drone pose, target state with visibility flag, bilingual (Chinese-English) instructions, and trajectory-pair metadata. To generate high-quality expert trajectories, we develop MuCO, a multi-constraint optimizer that plans directly in continuous three-dimensional space with BVH-accelerated collision and visibility queries, jointly enforcing target visibility, viewpoint quality, collision avoidance, smoothness, and kinematic feasibility, avoiding the discretization artifacts and post-hoc smoothing of grid-based planners. Fine-tuning experiments on seven vision-language models show that CosFlyTrack improves tracking performance to 78.3 to 95.6 percent SR@1 meter, a 53 to 69 percentage point gain over zero-shot baselines, supporting the dataset as a training resource for dynamic target-following agents. The dataset is publicly available at https://huggingface.co/datasets/AutelRobotics/CosFly; evaluation scripts and pre-trained checkpoints are hosted at https://huggingface.co/AutelRobotics/CosFly-Track.

无人机追踪多模态数据集视觉语言导航轨迹优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。