揭示AID型双层优化的泛化能力,理论证明其稳定性与单层非凸优化相当。
Exploring the Generalization Capabilities of AID-based Bi-level Optimization
- 基于统一稳定性分析,证明AID方法在非凸外层下仍具稳定收敛性
- 设计合适步长保证收敛,理论推导出与单层优化相当的泛化误差界
- 实验证实参数敏感性低,适用于真实任务场景
双层优化在现代机器学习中取得显著成功,尤其在超参数合理时。现有研究主要分为基于近似隐式微分(AID)和基于迭代微分(ITD)的方法。ITD方法可转化为单层优化,便于泛化分析;而AID方法因需保持双层结构,其泛化性质长期不明。本文在外部目标函数非凸条件下,证明了AID方法具有统一稳定性,其表现接近单层非凸问题。通过精心设计步长实现收敛性分析,结合稳定性结果,给出了AID方法的泛化能力。进一步开展参数消融实验,评估其在真实任务中的性能。实验结果验证了理论发现,表明该方法有效且具备广泛应用潜力。
原文摘要 · Abstract (English)
Bi-level optimization has achieved considerable success in contemporary machine learning applications, especially for given proper hyperparameters. However, due to the two-level optimization structure, commonly, researchers focus on two types of bi-level optimization methods: approximate implicit differentiation (AID)-based and iterative differentiation (ITD)-based approaches. ITD-based methods can be readily transformed into single-level optimization problems, facilitating the study of their generalization capabilities. In contrast, AID-based methods cannot be easily transformed similarly but must stay in the two-level structure, leaving their generalization properties enigmatic. In this paper, although the outer-level function is nonconvex, we ascertain the uniform stability of AID-based methods, which achieves similar results to a single-level nonconvex problem. We conduct a convergence analysis for a carefully chosen step size to maintain stability. Combining the convergence and stability results, we give the generalization ability of AID-based bi-level optimization methods. Furthermore, we carry out an ablation study of the parameters and assess the performance of these methods on real-world tasks. Our experimental results corroborate the theoretical findings, demonstrating the effectiveness and potential applications of these methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。