用全局统计信息提升遥感目标检索的判别能力
GLRT-Based Metric Learning for Remote Sensing Object Retrieval
- 基于广义似然比检验,融合全局数据分布优化样本难易评估
- 在FGSRSI-23和MAR20数据集上提升检索准确率1.8%~4.3%
- 适合处理训练与测试域分布偏移的遥感图像检索任务
随着遥感图像数量与质量的提升,基于内容的遥感目标检索(CBRSOR)日益重要。现有方法在训练与测试阶段忽略全局统计信息,导致网络过拟合于简单样本对,度量性能不佳。受奈曼-皮尔逊定理启发,本文提出广义似然比检验度量学习(GLRTML),通过引入全局数据分布信息,在训练与测试阶段估计样本对相对难度,引导网络关注困难样本,从而学习更具判别性的特征表示。此外,相较于传统度量空间,GLRTML因利用全局分布信息而更有效。准确估计嵌入分布是关键,但在实际应用中常存在训练域与目标域分布偏移,直接影响性能。为此,提出聚类伪标签快速参数适配(CPLFPA)方法,通过聚类目标域样本并重估分布参数,高效适应目标域分布。基于细粒度舰船遥感图像切片(FGSRSI-23)与军用飞机识别(MAR20)数据集重构实验,大量实验证明了GLRTML与CPLFPA的有效性。
原文摘要 · Abstract (English)
With the improvement in the quantity and quality of remote sensing images, content-based remote sensing object retrieval (CBRSOR) has become an increasingly important topic. However, existing CBRSOR methods neglect the utilization of global statistical information during both training and test stages, which leads to the overfitting of neural networks to simple sample pairs of samples during training and suboptimal metric performance. Inspired by the Neyman-Pearson theorem, we propose a generalized likelihood ratio test-based metric learning (GLRTML) approach, which can estimate the relative difficulty of sample pairs by incorporating global data distribution information during training and test phases. This guides the network to focus more on difficult samples during the training process, thereby encourages the network to learn more discriminative feature embeddings. In addition, GLRT is a more effective than traditional metric space due to the utilization of global data distribution information. Accurately estimating the distribution of embeddings is critical for GLRTML. However, in real-world applications, there is often a distribution shift between the training and target domains, which diminishes the effectiveness of directly using the distribution estimated on training data. To address this issue, we propose the clustering pseudo-labels-based fast parameter adaptation (CPLFPA) method. CPLFPA efficiently estimates the distribution of embeddings in the target domain by clustering target domain instances and re-estimating the distribution parameters for GLRTML. We reorganize datasets for CBRSOR tasks based on fine-grained ship remote sensing image slices (FGSRSI-23) and military aircraft recognition (MAR20) datasets. Extensive experiments on these datasets demonstrate the effectiveness of our proposed GLRTML and CPLFPA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。