揭示掩码预训练的理论机制,提出改进方案提升模型性能
A Theoretical Framework for Masked Pretraining (MPT)

- 从理论上阐明掩码如何隐式生成语义相似正样本
- 发现掩码导致维度坍塌问题,并设计新损失函数解决
- 提出新型掩码策略,适用于图像、文本等多领域下游任务
基于重建任务的掩码预训练(MPT)已成为跨领域的自监督学习主流范式,在多个下游任务中表现优异。然而其内在工作机制的理论理解仍不充分。本文建立新的理论框架,揭示掩码通过隐式构造语义相似正样本,使重建损失将它们拉近特征空间。同时指出该机制引发维度坍塌问题,提出增强均匀性的U-MPT损失,显著提升线性评估、跨数据集微调及分布外泛化能力。进一步建立U-MPT的下游保证,分析不同掩码策略的影响,并据此提出新掩码策略,从理论上解释现有方法的改进效果。
原文摘要 · Abstract (English)
Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. However, the theoretical understanding of the working mechanism behind MPT is still limited. In this paper, we introduce a new theoretical framework to analyze MPT and understand the crucial role of masking in extracting meaningful representations. We establish theoretical connections between MPT and another popular self-supervised paradigm: contrastive learning. We prove that the masking technique implicitly creates positive pairs that are semantically similar and the reconstruction loss pulls them together in the feature space. Besides, as a result of the implicit alignment, we point out the dimensional collapse issue of MPT and propose a Uniformity-enhanced MPT (U-MPT) loss that can effectively address this issue and bring significant improvements in downstream tasks including linear evaluation, cross-dataset fine-tuning and out-of-distribution generalization on real-world data sets. Furthermore, we establish downstream guarantees of U-MPT and theoretically analyze the influence of masking strategies. Based on the theoretical analysis, we propose a new masking strategy which enhances the downstream performance of MPT and explains current improvements of masking strategies with our theoretical perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。