arXiv:2607.08075cs.CV2026-07

为无人机视频设计开放词汇实例分割新任务,实现灵活目标查询与精准追踪

UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation

论文配图:UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation
图 1 · 摘自论文原文
  • 无需训练,用现有视觉模型协同完成目标发现与分割
  • 在8279条轨迹上表现优于迁移的现有方法,支持长视频密集目标处理
  • 适合做无人机感知、城市监控等开放场景下的实例理解研究

无人机视频广泛应用于交通监控、城市管理与应急救援。但现有感知方法多局限于预定义类别的框级检测与跟踪,难以支持开放场景下的灵活查询和细粒度时序理解。为此,我们提出新任务——无人机开放词汇视频实例分割(UAV-OVVIS),旨在根据开放词汇查询发现目标,并输出具有全局一致身份的实例分割轨迹。针对无人机场景实例标注稀缺问题,提出AeroTrack框架,通过周期性开放词汇检测与短片段掩码传播实现目标发现与分割,并引入生命周期感知身份关联(LIA)恢复分段推理下的全局身份。基于该框架,构建了包含9类物体和8,279条轨迹的AeroVIS基准数据集。实验表明,AeroTrack在性能上优于迁移的现有开放词汇视频实例分割方法,且在长视频中展现出良好的开放词汇迁移能力与密集目标处理能力。代码与数据集将在论文接收后开源。

原文摘要 · Abstract (English)

Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video perception is largely limited to box-level detection and tracking over predefined categories, making it difficult to jointly support flexible queries and fine-grained instance-level understanding of temporal dynamics in open scenarios. To this end, we introduce a new task, UAV Open-Vocabulary Video Instance Segmentation (UAV-OVVIS), which aims to discover targets in UAV videos according to open-vocabulary queries and output instance segmentation trajectories with globally consistent identities. Considering the scarcity of instance-level annotations in UAV scenarios, we propose AeroTrack, a training-free framework that coordinates existing visual foundation models to realize UAV-OVVIS. AeroTrack performs target discovery and segmentation through periodic open-vocabulary detection and short-segment mask propagation, and introduces Lifecycle-aware ID Association (LIA) to recover global identities under segment-wise inference. Based on this framework, we instantiate five feasible variants and construct AeroVIS, a UAV-OVVIS evaluation benchmark containing 9 UAV object categories and 8,279 trajectories. Experiments show that AeroTrack achieves better overall performance than the evaluated OV-VIS methods transferred to AeroVIS, while demonstrating good open-vocabulary transferability and dense-target handling capability in long UAV videos. The AeroTrack framework and the AeroVIS dataset will be open-sourced upon acceptance.

无人机视觉开放词汇实例分割视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。