arXiv:2604.02773cs.CV2026-04被引 2

提出点提示小目标检测新范式,显著提升弱标注下小目标识别效果。

Generalized Small Object Detection:A Point-Prompted Paradigm and Benchmark

  • 引入推理时稀疏点提示,实现类别级定位的语义增强。
  • 在TinySet-9M上仅用单次点击,相对全监督基线提升31.4%(AP75)。
  • 构建首个大规模多领域小目标数据集,支持可迁移检测框架训练。

小目标检测因像素极少、边界模糊而困难,导致标注难、高质量数据稀缺,且语义表征薄弱。本文首先构建了首个大规模多领域小目标检测数据集TinySet-9M,填补了大尺度数据空白,并建立了评估标签高效检测方法的基准。评估发现,弱视觉线索进一步加剧标签高效方法在小目标检测中的性能退化,凸显关键挑战。为此,提出推理时点提示的小目标检测新范式(P2SOD),通过稀疏点提示作为类别级定位的信息桥梁,实现语义增强。基于该范式与TinySet-9M,开发出可扩展、可迁移的点提示检测框架DEAL。仅需一次点击,DEAL在TinySet-9M上以严格定位指标(如AP75)相较全监督基线实现31.4%相对提升,并有效泛化至未见类别和未见数据集。

原文摘要 · Abstract (English)

Small object detection (SOD) remains challenging due to extremely limited pixels and ambiguous object boundaries. These characteristics lead to challenging annotation, limited availability of large-scale high-quality datasets, and inherently weak semantic representations for small objects. In this work, we first address the data limitation by introducing TinySet-9M, the first large-scale, multi-domain dataset for small object detection. Beyond filling the gap in large-scale datasets, we establish a benchmark to evaluate the effectiveness of existing label-efficient detection methods for small objects. Our evaluation reveals that weak visual cues further exacerbate the performance degradation of label-efficient methods in small object detection, highlighting a critical challenge in label-efficient SOD. Secondly, to tackle the limitation of insufficient semantic representation, we move beyond training-time feature enhancement and propose a new paradigm termed Point-Prompt Small Object Detection (P2SOD). This paradigm introduces sparse point prompts at inference time as an efficient information bridge for category-level localization, enabling semantic augmentation. Building upon the P2SOD paradigm and the large-scale TinySet-9M dataset, we further develop DEAL (DEtect Any smalL object), a scalable and transferable point-prompted detection framework that learns robust, prompt-conditioned representations from large-scale data. With only a single click at inference time, DEAL achieves a 31.4% relative improvement over fully supervised baselines under strict localization metrics (e.g., AP75) on TinySet-9M, while generalizing effectively to unseen categories and unseen datasets. Our project is available at https://zhuhaoraneis.github.io/TinySet-9M/.

小目标检测点提示数据集弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。