提升红外图像自编码器性能,通过双域引导增强关键信息建模。
DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain Guidance

- 基于令牌熵设计确定性掩码,保留高熵信息以提升重建质量。
- 引入双域引导模块,同时建模全局关系并自适应过滤非均匀噪声。
- 在包含59万张图像的Inf-590K数据集上预训练,适用于多种红外任务。
红外成像在低光照和恶劣天气条件下至关重要。然而,由于红外图像的独特特性,现有在可见光数据上训练的基础模型(如掩码自编码器MAE)在红外图像理解任务中表现不佳。为此,研究人员开发了基于红外数据的大规模预训练模型InfMAE。尽管有效,InfMAE仍存在忽略有用令牌、全局关联建模不足及忽视非均匀噪声等问题。本文提出基于MAE的双域引导红外基础模型DuGI-MAE:首先设计基于令牌熵的确定性掩码策略,仅保留高熵令牌用于重建,以增强信息量;其次引入双域引导(DDG)模块,同时捕捉全局令牌关系并自适应滤除红外图像中常见的非均匀背景噪声。为支持大规模预训练,构建了包含多样化场景、目标类型与空间分辨率的Inf-590K红外图像数据集。在该数据集上预训练后,DuGI-MAE在红外目标检测、语义分割和小目标检测等下游任务中展现出优异的泛化能力。实验结果表明,其性能优于多种监督与自监督方法。代码已提供于附录。
原文摘要 · Abstract (English)
Infrared imaging plays a critical role in low-light and adverse weather conditions. However, due to the distinct characteristics of infrared images, existing foundation models such as Masked Autoencoder (MAE) trained on visible data perform suboptimal in infrared image interpretation tasks. To bridge this gap, an infrared foundation model known as InfMAE was developed and pre-trained on large-scale infrared datasets. Despite its effectiveness, InfMAE still faces several limitations, including the omission of informative tokens, insufficient modeling of global associations, and neglect of non-uniform noise. In this paper, we propose a Dual-domain Guided Infrared foundation model based on MAE (DuGI-MAE). First, we design a deterministic masking strategy based on token entropy, preserving only high-entropy tokens for reconstruction to enhance informativeness. Next, we introduce a Dual-Domain Guidance (DDG) module, which simultaneously captures global token relationships and adaptively filters non-uniform background noise commonly present in infrared imagery. To facilitate large-scale pretraining, we construct Inf-590K, a comprehensive infrared image dataset encompassing diverse scenes, various target types, and multiple spatial resolutions. Pretrained on Inf-590K, DuGI-MAE demonstrates strong generalization capabilities across various downstream tasks, including infrared object detection, semantic segmentation, and small target detection. Experimental results validate the superiority of the proposed method over both supervised and self-supervised comparison methods. Our code is available in the supplementary material.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。