arXiv:2502.07216cs.CVcs.AI2025-02中稿 · ACM MM 2024被引 8

针对超高清广角图像中的稀疏目标,提出高效检测模型SparseFormer

SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer

  • 用稀疏注意力机制聚焦可能含目标的区域,兼顾全局与局部特征
  • 在PANDA和DOTA-v1.0上提升精度最多5.8%,速度最快达3倍
  • 适合处理高分辨率广角图像中的小目标检测任务

近年来,大分辨率(吉像素级)图像与视频采集系统及高分辨率广角(HRW)基准数据集日益普及。然而,与MS COCO数据集中近距离拍摄不同,更高分辨率和更宽视野带来极端稀疏性和巨大尺度变化,导致现有近距检测器精度低、效率差。本文提出一种新型模型无关的稀疏视觉变换器SparseFormer,通过选择性使用注意力令牌,精准分析稀疏分布的窗口以定位潜在目标。该方法融合粗粒度与细粒度特征,联合实现全局与局部注意力,有效应对尺度剧烈变化。SparseFormer还引入一种新型跨切片非极大值抑制(C-NMS)算法,从噪声窗口中精确定位目标,并采用简单高效的多尺度策略提升精度。在两个HRW基准数据集PANDA和DOTA-v1.0上的大量实验表明,该模型相比当前最优方法,检测精度最高提升5.8%,速度最快提升3倍。

原文摘要 · Abstract (English)

Recent years have seen an increase in the use of gigapixel-level image and video capture systems and benchmarks with high-resolution wide (HRW) shots. However, unlike close-up shots in the MS COCO dataset, the higher resolution and wider field of view raise unique challenges, such as extreme sparsity and huge scale changes, causing existing close-up detectors inaccuracy and inefficiency. In this paper, we present a novel model-agnostic sparse vision transformer, dubbed SparseFormer, to bridge the gap of object detection between close-up and HRW shots. The proposed SparseFormer selectively uses attentive tokens to scrutinize the sparsely distributed windows that may contain objects. In this way, it can jointly explore global and local attention by fusing coarse- and fine-grained features to handle huge scale changes. SparseFormer also benefits from a novel Cross-slice non-maximum suppression (C-NMS) algorithm to precisely localize objects from noisy windows and a simple yet effective multi-scale strategy to improve accuracy. Extensive experiments on two HRW benchmarks, PANDA and DOTA-v1.0, demonstrate that the proposed SparseFormer significantly improves detection accuracy (up to 5.8%) and speed (up to 3x) over the state-of-the-art approaches.

目标检测稀疏注意力高分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。