arXiv:2503.00782cs.CV2025-03AAAI被引 6

用小波变换加速图像自监督学习,提升训练效率。

Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual Representation

  • 用小波变换分层分解图像,生成多尺度重建目标。
  • 在多个下游任务上性能相当或更优,训练速度更快。
  • 适合追求高效视觉表征学习的研究者。

掩码图像建模(MIM)在自监督学习中备受关注,因其能为下游任务学习可扩展的视觉表征。然而,图像本身包含大量冗余信息,导致基于像素的MIM重建过程过度关注纹理等细节,不必要的延长了训练时间。解决该问题需在重建过程中采用紧凑的特征表示。频域分析为此提供了可行路径。与常用的傅里叶变换不同,小波变换不仅提供频率信息,还保留空间特性与多层级图像特征。此外,小波变换的多级分解过程与现代神经网络的层次结构高度契合。本研究利用小波变换作为高效表征学习工具,加速MIM训练。具体而言,通过小波变换对图像进行多级分解,使用不同层级的小波系数构建代表不同频率与尺度的重建目标,并将其融入MIM流程,通过可调权重优先处理关键信息。大量实验表明,该方法在多个下游任务上达到相当或更优性能,同时显著提升训练效率。

原文摘要 · Abstract (English)

Masked Image Modeling (MIM) has garnered significant attention in self-supervised learning, thanks to its impressive capacity to learn scalable visual representations tailored for downstream tasks. However, images inherently contain abundant redundant information, leading the pixel-based MIM reconstruction process to focus excessively on finer details such as textures, thus prolonging training times unnecessarily. Addressing this challenge requires a shift towards a compact representation of features during MIM reconstruction. Frequency domain analysis provides a promising avenue for achieving compact image feature representation. In contrast to the commonly used Fourier transform, wavelet transform not only offers frequency information but also preserves spatial characteristics and multi-level features of the image. Additionally, the multi-level decomposition process of wavelet transformation aligns well with the hierarchical architecture of modern neural networks. In this study, we leverage wavelet transform as a tool for efficient representation learning to expedite the training process of MIM. Specifically, we conduct multi-level decomposition of images using wavelet transform, utilizing wavelet coefficients from different levels to construct distinct reconstruction targets representing various frequencies and scales. These reconstruction targets are then integrated into the MIM process, with adjustable weights assigned to prioritize the most crucial information. Extensive experiments demonstrate that our method achieves comparable or superior performance across various downstream tasks while exhibiting higher training efficiency.

自监督学习小波变换图像建模高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。