arXiv:2507.11040cs.CV2025-07ICCV被引 4

用Transformer提升高分辨率卫星图像检测效率,性能领先。

Combining Transformers and CNNs for Efficient Object Detection in High-Resolution Satellite Imagery

  • 以Swin Transformer替代CNN主干,实现端到端特征提取
  • 在xView数据集上达32.95%精度,超越当前最优方法11.46%
  • 适合需要高效、精准检测的遥感图像应用

我们提出GLOD,一种面向高分辨率卫星影像的目标检测的Transformer优先架构。GLOD采用Swin Transformer取代传统CNN主干,实现端到端特征提取,并引入新型的UpConvMixer块进行鲁棒上采样,以及Fusion Blocks实现多尺度特征融合。该方法在xView数据集上取得32.95%的精度,较现有最优方法提升11.46%。关键创新包括结合CBAM注意力的非对称融合机制与多路径检测头设计,可有效捕捉不同尺度目标。架构针对卫星影像挑战优化,利用空间先验的同时保持计算效率。

原文摘要 · Abstract (English)

We present GLOD, a transformer-first architecture for object detection in high-resolution satellite imagery. GLOD replaces CNN backbones with a Swin Transformer for end-to-end feature extraction, combined with novel UpConvMixer blocks for robust upsampling and Fusion Blocks for multi-scale feature integration. Our approach achieves 32.95\% on xView, outperforming SOTA methods by 11.46\%. Key innovations include asymmetric fusion with CBAM attention and a multi-path head design capturing objects across scales. The architecture is optimized for satellite imagery challenges, leveraging spatial priors while maintaining computational efficiency.

目标检测卫星图像Transformer多尺度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。