arXiv:2606.02008stat.MLcs.LG2026-06

提出复杂度最小化框架,解释预训练数据越多下游越高效

Provable Data Scaling Law for Meta Learning via Complexity Minimization

  • 用下游模型复杂度最小化来学习元表示
  • 理论证明数据量增大时少样本适应误差下降
  • 适用于提升元学习的样本效率,适合研究者参考

预训练已成为现代机器学习的基础范式,其关键经验优势在于:随着预训练数据规模增加,下游任务所需样本量减少。然而,现有预训练理论框架未能充分解释此现象。本文提出复杂度最小化这一新型元表示学习框架,通过评估每个领域最适配的下游模型复杂度,并在源域间最小化最坏情况下的复杂度,实现理论分析。我们完成了从预训练到下游回归的端到端理论分析,证明该框架能严格捕捉这一缩放规律;特别地,证明了少样本适应的误差率随元训练数据量增加而降低。实验上,将复杂度正则化引入现有元学习方法,可持续提升下游样本效率。

原文摘要 · Abstract (English)

Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-training data increases. However, existing theoretical frameworks for pre-training do not fully explain this phenomenon. In this paper, we introduce complexity minimization, a novel meta-representation learning framework designed to enable theoretical analysis of this scaling behavior, which learns representations by evaluating the downstream model complexity best suited to each domain and minimizing the worst-case such complexity across source domains. Our end-to-end theoretical analysis, spanning pre-training through downstream regression, shows that this framework provably captures this scaling behavior; in particular, we show that the error rate of few-shot adaptation improves as the amount of meta-training data grows. Empirically, we demonstrate that incorporating complexity regularization into existing meta-learning methods consistently improves downstream sample efficiency.

元学习复杂度预训练理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。