arXiv:2602.03531cs.LGcs.CV2026-02被引 1

MAE模型学习到的特征对模糊、遮挡等退化具有强鲁棒性。

Robust Representation Learning in Masked Autoencoders

  • 通过层间分析发现MAE逐步构建类感知的潜在空间。
  • 在模糊和遮挡下仍保持高分类性能,关键特征保留率超85%。
  • 首次量化了特征鲁棒性,适合研究视觉表征与鲁棒性。

掩码自编码器(MAE)在图像分类任务中表现优异,但其内部表征机制仍不清晰。本文旨在理解MAE为何具备强大下游分类能力。研究发现,预训练与微调后的表征具有显著鲁棒性,在模糊、遮挡等退化条件下仍能保持良好分类性能。通过逐层分析令牌嵌入,我们发现预训练的MAE在网络深度上逐步构建类感知的潜在空间:不同类别的嵌入逐渐分布于可分离的子空间中。此外,我们观察到MAE在编码器各层表现出早期且持续的全局注意力,有别于标准视觉变压器(ViT)。为量化特征鲁棒性,我们引入两个敏感性指标:干净与扰动嵌入间的方向一致性,以及退化下各头活性特征的保留率。这些研究揭示了MAE强大分类性能的内在原因。

原文摘要 · Abstract (English)

Masked Autoencoders (MAEs) achieve impressive performance in image classification tasks, yet the internal representations they learn remain less understood. This work started as an attempt to understand the strong downstream classification performance of MAE. In this process we discover that representations learned with the pretraining and fine-tuning, are quite robust -- demonstrating a good classification performance in the presence of degradations, such as blur and occlusions. Through layer-wise analysis of token embeddings, we show that pretrained MAE progressively constructs its latent space in a class-aware manner across network depth: embeddings from different classes lie in subspaces that become increasingly separable. We further observe that MAE exhibits early and persistent global attention across encoder layers, in contrast to standard Vision Transformers (ViTs). To quantify feature robustness, we introduce two sensitivity indicators: directional alignment between clean and perturbed embeddings, and head-wise retention of active features under degradations. These studies help establish the robust classification performance of MAEs.

鲁棒性表征学习自编码器视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。