无需标注,让雷达图像中的目标与背景自动分离。
Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
- 融合高层语义与低层散射特征,构建多层级表示
- 通过注意力机制提升目标槽的区分度,性能领先现有方法
- 适合雷达图像目标识别、无需人工标注的场景
合成孔径雷达(SAR)图像包含感兴趣目标及复杂背景杂波,如地表反射和斑点噪声。在许多情况下,这些杂波的强度和模式与目标相似,导致模型提取纠缠或虚假特征,削弱目标表征能力。为此,我们提出一种新型无掩码标注的目标中心学习框架SlotSAR。该方法首先从SARATR-X提取高层语义特征,从小波散射网络获取低层散射特征,以获得互补的多层级表示,用于鲁棒目标表征。进一步设计多层级槽注意力模块,融合高低层特征,增强槽级表示的区分性,实现有效的目标中心学习。实验表明,相比现有目标中心学习方法,SlotSAR在保留结构细节方面达到最新性能。
原文摘要 · Abstract (English)
Synthetic aperture radar (SAR) images contain not only targets of interest but also complex background clutter, including terrain reflections and speckle noise. In many cases, such clutter exhibits intensity and patterns that resemble targets, leading models to extract entangled or spurious features. Such behavior undermines the ability to form clear target representations, regardless of the classifier. To address this challenge, we propose a novel object-centric learning (OCL) framework, named SlotSAR, that disentangles target representations from background clutter in SAR images without mask annotations. SlotSAR first extracts high-level semantic features from SARATR-X and low-level scattering features from the wavelet scattering network in order to obtain complementary multi-level representations for robust target characterization. We further present a multi-level slot attention module that integrates these low- and high-level features to enhance slot-wise representation distinctiveness, enabling effective OCL. Experimental results demonstrate that SlotSAR achieves state-of-the-art performance in SAR imagery by preserving structural details compared to existing OCL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。