用密度图引导注意力,提升遥感图像中密集微小目标检测效果
Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- 通过密度图作为空间先验,动态引导网络聚焦密集区域
- 在AI-TOD和DTOD数据集上达到最新最优性能,尤其在高密度遮挡场景下
- 适合遥感图像分析、无人机巡检等需要精准检测微小目标的场景
高分辨率遥感影像中密集微小目标检测面临严重相互遮挡和像素覆盖范围有限的挑战。现有方法通常均匀分配计算资源,难以自适应聚焦于高密度区域,影响特征学习效果。为此,我们提出密集区域挖掘网络(DRMNet),利用密度图作为显式空间先验,引导自适应特征学习。首先,设计密度生成分支(DGB)建模物体分布模式,提供可量化的先验信息以引导网络关注密集区域。其次,为解决全局注意力的计算瓶颈,提出密集区域聚焦模块(DAFM),基于密度图识别并聚焦密集区域,实现高效局部-全局特征交互。最后,为缓解层级提取中的特征退化问题,引入双滤波融合模块(DFFM),通过离散余弦变换将多尺度特征解耦为高低频成分,并进行密度引导的跨注意力增强互补性,抑制背景干扰。在AI-TOD和DTOD数据集上的大量实验表明,DRMNet优于当前最先进方法,尤其在高密度与严重遮挡复杂场景中表现突出。
原文摘要 · Abstract (English)
High-resolution remote sensing imagery increasingly contains dense clusters of tiny objects, the detection of which is extremely challenging due to severe mutual occlusion and limited pixel footprints. Existing detection methods typically allocate computational resources uniformly, failing to adaptively focus on these density-concentrated regions, which hinders feature learning effectiveness. To address these limitations, we propose the Dense Region Mining Network (DRMNet), which leverages density maps as explicit spatial priors to guide adaptive feature learning. First, we design a Density Generation Branch (DGB) to model object distribution patterns, providing quantifiable priors that guide the network toward dense regions. Second, to address the computational bottleneck of global attention, our Dense Area Focusing Module (DAFM) uses these density maps to identify and focus on dense areas, enabling efficient local-global feature interaction. Finally, to mitigate feature degradation during hierarchical extraction, we introduce a Dual Filter Fusion Module (DFFM). It disentangles multi-scale features into high- and low-frequency components using a discrete cosine transform and then performs density-guided cross-attention to enhance complementarity while suppressing background interference. Extensive experiments on the AI-TOD and DTOD datasets demonstrate that DRMNet surpasses state-of-the-art methods, particularly in complex scenarios with high object density and severe occlusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。