arXiv:2409.08885cs.CV2024-09被引 1

提出交互式图像掩码建模,提升遥感图像小目标检测性能

Interactive Masked Image Modeling for Multimodal Object Detection in Remote Sensing

  • 设计交互式掩码建模,增强图像块间语义关联
  • 在遥感数据集上显著提升小目标检测准确率
  • 适合遥感图像分析与多模态学习研究者

遥感图像中的目标检测在地球观测应用中至关重要。然而,与自然场景图像相比,该任务因多样地形中存在大量微小且难辨识的目标而极具挑战性。多模态学习可通过融合不同数据模态特征提升检测精度,但其性能常受限于标注数据集规模较小。本文提出使用掩码图像建模(Masked Image Modeling, MIM)作为预训练技术,利用无标签数据进行自监督学习以增强检测表现。传统MIM方法(如MAE)仅对掩码块进行独立处理,缺乏与其他图像部分的交互,难以捕捉细粒度信息。为此,我们提出一种新型交互式MIM方法,可在不同图像块间建立交互,特别有利于遥感图像中的目标检测。大量消融实验与评估验证了该方法的有效性。

原文摘要 · Abstract (English)

Object detection in remote sensing imagery plays a vital role in various Earth observation applications. However, unlike object detection in natural scene images, this task is particularly challenging due to the abundance of small, often barely visible objects across diverse terrains. To address these challenges, multimodal learning can be used to integrate features from different data modalities, thereby improving detection accuracy. Nonetheless, the performance of multimodal learning is often constrained by the limited size of labeled datasets. In this paper, we propose to use Masked Image Modeling (MIM) as a pre-training technique, leveraging self-supervised learning on unlabeled data to enhance detection performance. However, conventional MIM such as MAE which uses masked tokens without any contextual information, struggles to capture the fine-grained details due to a lack of interactions with other parts of image. To address this, we propose a new interactive MIM method that can establish interactions between different tokens, which is particularly beneficial for object detection in remote sensing. The extensive ablation studies and evluation demonstrate the effectiveness of our approach.

遥感检测图像建模自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。