arXiv:2606.25312cs.CV2026-06

构建百万级多类别遥感检测数据集,训练通用检测模型。

LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

论文配图:LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection
图 1 · 摘自论文原文
  • 设计分层感知的检测基础模型,支持跨传感器、多分辨率检测。
  • 在9个外部数据集上零样本超越现有方法,平均提升5.02 mAP。
  • 适合遥感目标检测、跨域泛化研究者使用。

遥感目标检测因大规模基准和现代检测架构的发展而迅速进步,但现有数据集与检测器仍呈碎片化。多数基准仅关注有限类别、固定空间分辨率或单一传感器,而检测器在跨传感器和类别体系时仍表现不佳。本文提出目前最大最全面的遥感目标检测数据集 LEVIRDet-159,包含159个类别、256万边界框及70万细粒度标注,采用多层级分类体系。在关键尺度维度上,其图像数量是现有最大数据集的7倍,目标实例数为6倍,类别数为4倍。基于该数据集,我们设计了面向通用遥感目标检测的尺度层次感知检测基础模型 LEVIRDetNet。该模型结合在线地表采样距离(GSD)预测、GSD条件查询调制与分配、以及层次感知检测头,实现混合粒度遥感监督。在严格评估设置下,即使未进行目标域训练或微调,其在9个外部基准上均达到领先性能,平均比最强的全监督对比方法提升5.02 mAP。我们希望本工作能推动跨多样化类别体系、空间分辨率与传感器平台的强泛化遥感目标检测发展。数据集与训练模型将公开于 https://qinzheyang.github.io/LEVIRDet/,随论文最终发布。

原文摘要 · Abstract (English)

Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existing datasets and detectors remain fragmented. Most benchmarks focus on limited categories, fixed spatial resolutions, or a single sensor, while detectors still struggle to work across different sensors and categorical systems. In this paper, we introduce LEVIRDet-159, the largest and most comprehensive remote sensing object detection dataset to date, with 159 categories, 2.56 million bounding boxes, and 700k fine-grained annotations under a multi-level taxonomy. In each key scale dimension, LEVIRDet-159 exceeds the corresponding largest existing remote sensing object detection dataset, containing approximately (7x) more images, (6x) more object instances, and (4x) more categories. Based on this dataset, we design LEVIRDetNet, a scale-hierarchy-aware detection foundation model for universal remote sensing object detection. LEVIRDetNet couples online visual Ground Sampling Distance (GSD) prediction, GSD-conditioned query modulation and allocation, and a hierarchy-aware detection head for mixed-granularity remote sensing supervision. Under stringent evaluation settings, LEVIRDetNet demonstrates strong cross-domain generalization. Even without target-domain training or fine-tuning, it achieves state-of-the-art detection performance on 9 external benchmarks, improving the strongest fully supervised competing methods by 5.02 mAP on average under each benchmark's primary metric. We hope this study will facilitate the development of strongly generalizable remote sensing object detection across diverse category systems, spatial resolutions, and sensor platforms. The dataset and trained models will be released at https://qinzheyang.github.io/LEVIRDet/, accompanying the final paper.

遥感检测多类别基础模型跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。