提出分布式极小极大优化的泛化分析框架,揭示泛化与优化的权衡关系。
Stability and Generalization for Distributed SGDA
- 基于稳定性理论构建统一分析框架,覆盖局部SGDA与去中心化SGDA
- 发现泛化误差与优化误差存在权衡,可指导超参数选择以最小化总体风险
- 适用于关注隐私保护和通信效率的分布式学习场景
极小极大优化在现代机器学习中日益受到重视。随着大规模模型和来自边缘设备的海量数据,以及对客户端隐私保护的关注,通信高效的分布式极小极大优化算法(如局部随机梯度下降上升,Local-SGDA,和局部去中心化SGDA,Local-DSGDA)变得流行。尽管现有研究多集中于收敛速度、计算复杂度和通信效率,但泛化性能仍缺乏系统研究,而泛化能力是评估模型在未知数据上表现的关键指标。本文为分布式SGDA提出基于稳定性的泛化分析框架,统一分析两种主流算法,并在不同设置下(如(S)C-(S)C、PL-SC、NC-NC)全面研究稳定性误差、泛化差距和种群风险。理论结果揭示了泛化差距与优化误差之间的权衡,建议超参数选择以获得最优种群风险。针对Local-SGDA和Local-DSGDA的数值实验验证了理论结论。
原文摘要 · Abstract (English)
Minimax optimization is gaining increasing attention in modern machine learning applications. Driven by large-scale models and massive volumes of data collected from edge devices, as well as the concern to preserve client privacy, communication-efficient distributed minimax optimization algorithms become popular, such as Local Stochastic Gradient Descent Ascent (Local-SGDA), and Local Decentralized SGDA (Local-DSGDA). While most existing research on distributed minimax algorithms focuses on convergence rates, computation complexity, and communication efficiency, the generalization performance remains underdeveloped, whereas generalization ability is a pivotal indicator for evaluating the holistic performance of a model when fed with unknown data. In this paper, we propose the stability-based generalization analytical framework for Distributed-SGDA, which unifies two popular distributed minimax algorithms including Local-SGDA and Local-DSGDA, and conduct a comprehensive analysis of stability error, generalization gap, and population risk across different metrics under various settings, e.g., (S)C-(S)C, PL-SC, and NC-NC cases. Our theoretical results reveal the trade-off between the generalization gap and optimization error and suggest hyperparameters choice to obtain the optimal population risk. Numerical experiments for Local-SGDA and Local-DSGDA validate the theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。