arXiv:2604.04090cs.LGcs.AI2026-04IJCAI被引 10

揭示了随机双层优化的泛化能力与稳定性关系,给出更通用的理论保障。

Fine-grained Analysis of Stability and Generalization for Stochastic Bilevel Optimization

  • 通过平均参数稳定性建立泛化误差上界
  • 在三种场景下推导出单/双尺度梯度下降的稳定性上界
  • 无需重置内层参数,适用于更广目标函数

随机双层优化(SBO)近年来被广泛应用于超参数优化、元学习和强化学习等机器学习范式。尽管其计算行为已有诸多研究,但从统计学习理论视角看,SBO方法的泛化保证仍不清晰。本文对基于一阶梯度的双层优化方法进行系统性泛化分析:首先建立平均参数稳定性与泛化差距之间的定量联系;随后分别在非凸-非凸(NC-NC)、凸-凸(C-C)和强凸-强凸(SC-SC)三种设置下,推导出单时间尺度随机梯度下降(SGD)和双时间尺度SGD的平均参数稳定性上界。实验验证了理论结论。相比以往算法稳定性分析,本结果无需每轮重置内层参数,且可适用于更一般的目标函数。

原文摘要 · Abstract (English)

Stochastic bilevel optimization (SBO) has been integrated into many machine learning paradigms recently, including hyperparameter optimization, meta learning, and reinforcement learning. Along with the wide range of applications, there have been numerous studies on the computational behavior of SBO. However, the generalization guarantees of SBO methods are far less understood from the lens of statistical learning theory. In this paper, we provide a systematic generalization analysis of the first-order gradient-based bilevel optimization methods. Firstly, we establish the quantitative connections between the on-average argument stability and the generalization gap of SBO methods. Then, we derive the upper bounds of on-average argument stability for single-timescale stochastic gradient descent (SGD) and two-timescale SGD, where three settings (nonconvex-nonconvex (NC-NC), convex-convex (C-C), and strongly-convex-strongly-convex (SC-SC)) are considered respectively. Experimental analysis validates our theoretical findings. Compared with the previous algorithmic stability analysis, our results do not require reinitializing the inner-level parameters at each iteration and are applicable to more general objective functions.

双层优化泛化分析稳定性随机优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。