arXiv:2502.19718cs.CV2025-02ICLR被引 5

通过信息瓶颈原理提升掩码图像建模的性能

Learning Mask Invariant Mutual Information for Masked Image Modeling

  • 基于信息瓶颈理论,优化隐藏特征的相关与无关信息平衡
  • 在图像分类等任务上显著优于传统MAE模型
  • 为自监督学习提供可解释的新思路,适合研究者参考

掩码自编码器(MAEs)是计算机视觉中重要的自监督学习范式。尽管其在实践中表现优异,但其内在机制仍不明确。现有研究多通过对比学习和特征表示分析揭示其工作原理,但往往仅提供隐含见解。本文基于信息论中的信息瓶颈原则,提出新视角理解MAEs。理论分析表明,优化隐藏特征以平衡相关与无关信息是提升性能的关键。基于此,我们提出MI-MAE方法,通过最大化与输出之间的互信息、最小化与输入之间的互信息来优化模型。实验在标准基准上显示,该方法在图像分类、目标检测和语义分割任务中显著优于传统MAE模型。结果验证了理论框架的有效性,展示了信息瓶颈原则在构建更强大自监督模型中的实际优势。

原文摘要 · Abstract (English)

Masked autoencoders (MAEs) represent a prominent self-supervised learning paradigm in computer vision. Despite their empirical success, the underlying mechanisms of MAEs remain insufficiently understood. Recent studies have attempted to elucidate the functioning of MAEs through contrastive learning and feature representation analysis, yet these approaches often provide only implicit insights. In this paper, we propose a new perspective for understanding MAEs by leveraging the information bottleneck principle in information theory. Our theoretical analyses reveal that optimizing the latent features to balance relevant and irrelevant information is key to improving MAE performance. Building upon our proofs, we introduce MI-MAE, a novel method that optimizes MAEs through mutual information maximization and minimization. By enhancing latent features to retain maximal relevant information between them and the output, and minimizing irrelevant information between them and the input, our approach achieves better performance. Extensive experiments on standard benchmarks show that MI-MAE significantly outperforms MAE models in tasks such as image classification, object detection, and semantic segmentation. Our findings validate the theoretical framework and highlight the practical advantages of applying the information bottleneck principle to MAEs, offering deeper insights for developing more powerful self-supervised learning models.

自监督学习信息瓶颈掩码图像建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。