arXiv:2603.27631cs.LGstat.ML2026-03

揭示自监督预训练的渐近规律,解析表示对称性对下游性能的影响。

On the Asymptotics of Self-Supervised Pre-training: Two-Stage M-Estimation and Representation Symmetry

  • 通过两阶段M估计构建自监督预训练的渐近理论框架。
  • 证明下游测试风险的极限分布与表示轨道不变性密切相关。
  • 在谱预训练等场景中显著优于已有方法,理论更精准。

自监督预训练利用大规模无标签数据学习表示以供下游微调,已成为现代机器学习的核心。尽管已有理论工作开始分析该范式,但现有界仍无法回答当前速率是否紧致,以及是否准确捕捉了预训练与微调之间的复杂互动。本文通过两阶段M估计发展了预训练的渐近理论。关键挑战在于预训练估计量通常仅在群对称性下可识别,这在表示学习中普遍存在,需谨慎处理。我们借助黎曼几何工具研究预训练表示的内在参数,并通过轨道不变性将其与下游预测器关联,精确刻画了下游测试风险的极限分布。我们将主要结果应用于谱预训练、因子模型和高斯混合模型等案例,当适用时,在问题特定因素上相比先前工作取得显著改进。

原文摘要 · Abstract (English)

Self-supervised pre-training, where large corpora of unlabeled data are used to learn representations for downstream fine-tuning, has become a cornerstone of modern machine learning. While a growing body of theoretical work has begun to analyze this paradigm, existing bounds leave open the question of how sharp the current rates are, and whether they accurately capture the complex interaction between pre-training and fine-tuning. In this paper, we address this gap by developing an asymptotic theory of pre-training via two-stage M-estimation. A key challenge is that the pre-training estimator is often identifiable only up to a group symmetry, a feature common in representation learning that requires careful treatment. We address this issue using tools from Riemannian geometry to study the intrinsic parameters of the pre-training representation, which we link with the downstream predictor through a notion of orbit-invariance, precisely characterizing the limiting distribution of the downstream test risk. We apply our main result to several case studies, including spectral pre-training, factor models, and Gaussian mixture models, and obtain substantial improvements in problem-specific factors over prior art when applicable.

自监督学习渐近理论表示对称性两阶段估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。