用大模型辅助主动学习,降低遥感目标检测标注成本。
Foundation-Assisted Active Learning for Object Detection Annotation

- 结合大模型与检测器,双源评估样本不确定性。
- 冷启动阶段标注效率提升,低预算下表现更优。
- 自动修正框位置,减少人工修正工作量。
遥感目标检测的标注成本高,现有主动学习方法在目标检测场景中仍面临定位与分类不确定性耦合、冷启动阶段定位噪声严重以及高召回候选框导致的伪多样性等问题。为此,我们提出一种基于基础模型协同的主动学习与半自动标注框架,构建由UPN+SAM2生成的参考定位源(SA-source)和检测器预测源(OD-source)组成的双源机制,并提出基于大模型增强的双源不确定性估计,在冷启动阶段通过联合建模定位一致性与分类置信度提升样本选择质量。进一步提出面向目标的多样性采样,利用DINOv2特征与SAM2掩码构建对象级表示,提升样本覆盖范围同时抑制伪多样性。针对半自动标注阶段的几何噪声,设计双源框切换策略,将噪声检测框替换为匹配的精细化框,显著降低人工修正负担。在DIOR、HRSC2016、DOTAv2和FAIR1M数据集上的实验表明,该方法在多数标注预算下达到更优或相当的效果,尤其在低预算条件下表现出更强的冷启动样本效率。
原文摘要 · Abstract (English)
The annotation cost for remote sensing object detection is high, while existing active learning methods still face several challenges in object detection scenarios, including the coupling of localization and classification uncertainty, severe localization noise in the cold-start stage, and pseudo-diversity caused by high-recall candidate proposals. To address these issues, we propose a foundation-model-collaborative active learning and semi-automatic annotation framework for efficient construction of remote sensing object detection datasets. We build a dual-source mechanism consisting of a reference localization source (SA-source) based on UPN+SAM2 and a detector prediction source (OD-source), and further propose a Foundation-model-enhanced Dual-Source Uncertainty estimation to improve sample selection quality in the cold-start stage by jointly modeling localization consistency and classification confidence. Furthermore, we propose Object-Centric Diversity Sampling, which constructs object-level representations using DINOv2 features and SAM2 masks to improve sample coverage while suppressing pseudo-diversity. To address geometric noise in the semi-automatic annotation stage, we design Dual-Source Box Switching, which replaces noisy detector boxes with matched refined boxes from the SA-source, thereby reducing the manual burden of box refinement. Experiments on DIOR, HRSC2016, DOTAv2, and FAIR1M show that our method achieves superior or comparable results under most annotation budgets, with notably stronger cold-start sample efficiency in the low-budget regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。