arXiv:2607.09078cs.CVcs.RO2026-07

首个面向无人机主动目标检测的大规模真实数据集,解决视角受限难题。

Toward Active Object Detection for UAVs in the Wild: A Large-Scale Dataset, Benchmark and Method

论文配图:Toward Active Object Detection for UAVs in the Wild: A Large-Scale Dataset, Benchmark and Method
图 1 · 摘自论文原文
  • 构建多视角全景图像与局部目标切片的无人机-地面主动检测数据集
  • 揭示现有强化学习方法在跨场景泛化中的显著性能差距
  • 提出融合先验知识的JEPA世界模型,提升状态表征能力

目标检测是众多无人机应用的基础,但长期受遮挡和目标像素稀少等问题困扰。主动目标检测(AOD)通过主动视觉提供新范式,但基于无人机的AOD研究因缺乏高质量数据集和评估基准而进展缓慢。本文提出ATRNet-LUDO,首个大规模真实世界无人机-地面主动目标检测(UGAOD)数据集,包含12.1万张多视角全景多目标空中图像及121万张局部单目标切片,覆盖40种场景下的10类车辆目标。该数据集支持无人机智能体交互与主动观测策略学习的多样化训练与测试环境构建。基于此,我们建立全面的AOD策略学习评估基准。现有AOD策略多依赖深度强化学习(DRL),但泛化能力差。在该基准上的评估揭示了训练与测试性能间的显著泛化差距,凸显解决方案的迫切性。为此,我们采用联合嵌入预测架构(JEPA)构建世界模型以增强状态表示学习,并提出AOD-JEPA,融合AOD特定先验知识。大量实验验证其有效性与优越性。我们期望ATRNet-LUDO与基准能推动UGAOD领域发展。数据集与代码将发布于https://github.com/Leo000ooo/LUDO_dataset。

原文摘要 · Abstract (English)

Object detection is a fundamental component in numerous Unmanned Aerial Vehicle (UAV) applications, yet it has long been plagued by hindrances like occlusion or target pixel scarcity. Active Object Detection (AOD) provides a novel paradigm to address these challenges via active vision, while UAV-based AOD research remains scarce due to the lack of high-quality datasets and benchmarks for algorithm development and evaluation. To fill this gap, this paper presents ATRNet-LUDO, the first large-scale real-world dataset for UAV-Ground Active Object Detection (UGAOD). It contains 121,000 multi-view panoramic multi-target aerial images and 1.21 million local single-target slices, covering 10 vehicle targets across 40 scenarios. It enables the construction of diverse training and testing environments for UAV agent interaction and active observation policy learning. Based on this dataset, we establish a comprehensive evaluation benchmark for AOD policy learning methods. Most existing AOD policies rely on Deep Reinforcement Learning (DRL) but suffer from poor generalization. Evaluations on our benchmark reveal a significant generalization gap between training and testing performance, highlighting an urgent need for solutions. To this end, we leverage the Joint Embedding Predictive Architecture (JEPA) to construct a world model that enhances state representation learning, and propose AOD-JEPA by incorporating AOD-specific prior knowledge. Extensive experiments validate its effectiveness and superiority. We hope ATRNet-LUDO and the benchmark will advance research in the UGAOD field. The dataset and code are soon available at https://github.com/Leo000ooo/LUDO_dataset.

主动检测无人机数据集强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。