arXiv:2508.15404cs.CV2025-08ICCV被引 1

揭示掩码自编码器如何通过超参数控制学习图像空间相关性

From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations

  • 从线性模型推导出掩码率与图块大小对空间特征的选择机制
  • 非线性模型可捕捉数据集的长程空间相关性,超越二阶统计
  • 为新数据集配置超参数提供理论指导,减少调参成本

掩码自编码器(MAE)已成为视觉基础模型的强大预训练方法。尽管有效,但在应用于新数据集时仍需大量超参数调优(如掩码率、图块大小、编码器/解码器层数)。现有理论研究多关注注意力模式与分层潜在变量模型,但未深入探讨超参数与下游任务性能之间的联系。本文研究了MAE如何学习输入图像中的空间相关性。我们从理论上推导了线性MAE所学特征,表明掩码率和图块大小可用于选择捕获短程与长程空间相关性的特征。进一步扩展至非线性MAE,证明其表示能适应数据集中实际的空间相关性,超越二阶统计。最后,提出实践性超参数选择建议。

原文摘要 · Abstract (English)

Masked Autoencoders (MAEs) have emerged as a powerful pretraining technique for vision foundation models. Despite their effectiveness, they require extensive hyperparameter tuning (masking ratio, patch size, encoder/decoder layers) when applied to novel datasets. While prior theoretical works have analyzed MAEs in terms of their attention patterns and hierarchical latent variable models, the connection between MAE hyperparameters and performance on downstream tasks is relatively unexplored. This work investigates how MAEs learn spatial correlations in the input image. We analytically derive the features learned by a linear MAE and show that masking ratio and patch size can be used to select for features that capture short- and long-range spatial correlations. We extend this analysis to non-linear MAEs to show that MAE representations adapt to spatial correlations in the dataset, beyond second-order statistics. Finally, we discuss some insights on how to select MAE hyper-parameters in practice.

自编码器视觉表征超参数分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。