arXiv:2603.09070cs.RO2026-03

用互联网视频自动推断无人机3D轨迹和类型,无需人工标注

3D UAV Trajectory Estimation and Classification from Internet Videos via Language Model

  • 通过语言模型自动搜寻并筛选无人机视频,实现无监督数据获取
  • 零样本迁移在公开数据集上表现接近顶尖水平,数据越多效果越好
  • 适合反无人机系统研究者,尤其关注低成本数据构建的场景

可靠的无人机3D轨迹估计是反无人机系统的基本需求,但大规模精确标注的轨迹数据获取成本高昂。本文提出一种新框架,直接从互联网规模的无人机视频中推断3D轨迹与机型信息,无需人工标注。首先,采用语言驱动的数据采集,自主发现并收集无人机相关视频,并通过视觉-语言推理逐步筛选任务相关片段;其次,引入无训练的跨模态标签生成模块,推断3D轨迹假设与无人机类型线索;最后,设计物理约束优化过程,保证轨迹的时间平滑性与运动一致性。生成的视频片段与轨迹标注可直接用于下游反无人机任务。在公开的3D无人机基准数据集上进行零样本迁移实验,结果表明:随着在线视频数据量增加,目标数据集上的零样本性能持续提升,无需目标域训练。该方法表现接近当前最优水平,展现出强鲁棒性与实际应用潜力。代码与数据集将在论文录用后发布。

原文摘要 · Abstract (English)

Reliable 3D trajectory estimation of unmanned aerial vehicles (UAVs) is a fundamental requirement for anti-UAV systems, yet the acquisition of large-scale and accurately annotated trajectory data remains prohibitively expensive. In this work, we present a novel framework that derives UAV 3D trajectories and category information directly from Internet-scale UAV videos, without relying on manual annotations. First, language-driven data acquisition is employed to autonomously discover and collect UAV-related videos, while vision-language reasoning progressively filters task-relevant segments. Second, a training-free cross-modal label generation module is introduced to infer 3D trajectory hypotheses and UAV type cues. Third, a physics-informed refinement process is designed to impose temporal smoothness and kinematic consistency on the estimated trajectories. The resulting video clips and trajectory annotations can be readily utilized for downstream anti-UAV tasks. To assess effectiveness and generalization, we conduct zero-shot transfer experiments on a public, well-annotated 3D UAV benchmark. Results reveal a clear data scaling behavior: as the amount of online video data increases, zero-shot transfer performance on the target dataset improves consistently, without any target-domain training. The proposed method closely approaches the current state-of-the-art, highlighting its robustness and applicability to real-world anti-UAV scenarios. Code and datasets will be released upon acceptance.

无人机轨迹视觉语言模型零样本学习反无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。