提出统一信息论框架,更精准刻画元学习泛化能力。
A Unified Information-Theoretic Framework for Meta-Learning Generalization
- 用单步推导构建统一信息论框架,融合环境与任务级依赖。
- 新界紧致性更强,随任务数和每任务样本量呈理想缩放。
- 揭示噪声与迭代算法的泛化机制,适用于Reptile和MAML等方法。
近年来,信息论泛化界在分析元学习算法泛化能力方面受到越来越多关注。然而,现有研究局限于两步界,无法同时考虑环境层级和任务层级的依赖关系,难以提供对元泛化差距的精确刻画。本文通过单步推导,建立统一的信息论框架,得到的元泛化界以多种信息度量表示,在紧致性、随采样任务数及每任务样本量的缩放行为、计算可处理性方面均显著优于此前工作。此外,通过梯度协方差分析,为两类噪声型和迭代型元学习算法(如使用全部元训练数据的Reptile,或任务内分离训练测试数据的模型无关元学习(MAML))提供了新的理论洞察。数值结果验证了所导界在捕捉元学习泛化动态方面的有效性。
原文摘要 · Abstract (English)
In recent years, information-theoretic generalization bounds have gained increasing attention for analyzing the generalization capabilities of meta-learning algorithms. However, existing results are confined to two-step bounds, failing to provide a sharper characterization of the meta-generalization gap that simultaneously accounts for environment-level and task-level dependencies. This paper addresses this fundamental limitation by developing a unified information-theoretic framework using a single-step derivation. The resulting meta-generalization bounds, expressed in terms of diverse information measures, exhibit substantial advantages over previous work, particularly in terms of tightness, scaling behavior associated with sampled tasks and samples per task, and computational tractability. Furthermore, through gradient covariance analysis, we provide new theoretical insights into the generalization properties of two classes of noisy and iterative meta-learning algorithms, where the meta-learner uses either the entire meta-training data (e.g., Reptile), or separate training and test data within the task (e.g., model agnostic meta-learning (MAML)). Numerical results validate the effectiveness of the derived bounds in capturing the generalization dynamics of meta-learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。