针对复杂背景中小目标检测难题,提出多尺度注意力与全局关系建模框架。
Small Object Detection in Complex Backgrounds with Multi-Scale Attention and Global Relation Modeling
- 通过小波下采样保留细粒度结构,融合空域与频域特征
- 在高阶特征阶段捕获长程依赖,抑制背景噪声干扰
- 跨尺度注意力实现细节与语义高效融合,适合遥感等场景
复杂背景下小目标检测仍面临特征退化、语义表征弱和定位不准等问题,主要由下采样操作和背景干扰导致。现有检测框架多针对通用目标设计,未能显式处理小目标的特性,如结构线索有限、对定位误差敏感。本文提出一种专为小目标检测设计的多层级特征增强与全局关系建模框架。具体包括:引入残差哈尔小波下采样模块,联合利用空域卷积特征与频域表示以保留精细结构;采用全局关系建模模块,在高层特征阶段捕捉长程依赖,提升语义感知并抑制背景噪声;设计跨尺度混合注意力模块,在多尺度特征间建立稀疏对齐交互,实现高分辨率细节与高层语义的有效融合,且计算开销低;最后引入中心辅助损失,稳定训练并提升小目标定位精度。在大规模RGBT-Tiny数据集上的大量实验表明,所提方法在基于IoU和尺度自适应评价指标上均持续优于现有最先进检测器,验证了该框架在复杂环境中小目标检测中的有效性和鲁棒性。
原文摘要 · Abstract (English)
Small object detection under complex backgrounds remains a challenging task due to severe feature degradation, weak semantic representation, and inaccurate localization caused by downsampling operations and background interference. Existing detection frameworks are mainly designed for general objects and often fail to explicitly address the unique characteristics of small objects, such as limited structural cues and strong sensitivity to localization errors. In this paper, we propose a multi-level feature enhancement and global relation modeling framework tailored for small object detection. Specifically, a Residual Haar Wavelet Downsampling module is introduced to preserve fine-grained structural details by jointly exploiting spatial-domain convolutional features and frequency-domain representations. To enhance global semantic awareness and suppress background noise, a Global Relation Modeling module is employed to capture long-range dependencies at high-level feature stages. Furthermore, a Cross-Scale Hybrid Attention module is designed to establish sparse and aligned interactions across multi-scale features, enabling effective fusion of high-resolution details and high-level semantic information with reduced computational overhead. Finally, a Center-Assisted Loss is incorporated to stabilize training and improve localization accuracy for small objects. Extensive experiments conducted on the large-scale RGBT-Tiny benchmark demonstrate that the proposed method consistently outperforms existing state-of-the-art detectors under both IoU-based and scale-adaptive evaluation metrics. These results validate the effectiveness and robustness of the proposed framework for small object detection in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。