揭示了分布式学习中CoCoA与ADMM的深层联系,统一了两类算法的更新形式。
On the Relationship Between CoCoA and ADMM for Distributed Empirical Risk Minimization
- 从原对偶视角重构算法,发现两者可用统一框架表达
- 在岭回归下,调优后的ADMM性能不弱于甚至优于CoCoA
- 提供通用停止准则和收敛性分析,适合分布式优化研究者
分布式经验风险最小化常通过两大类方法研究:源于对偶坐标上升的CoCoA型算法,以及源于一致性与近端分裂的ADMM型算法。本文从统一的原对偶视角探讨二者关系,证明共识ADMM、线性化共识ADMM、两种分布式近端ADMM变体及岭正则化CoCoA均可写成包含全局原变量与块对偶变量的统一更新形式。该重构显式揭示了隐藏关联:在岭正则化ERM下,CoCoA与特定近端ADMM方案在对偶更新层面完全一致;共识ADMM在原问题上等价于对偶问题上的近端ADMM,仅需参数映射与鞍点目标符号反转,线性化情形亦同。结果表明,经调优的ADMM在岭正则化问题上至少不劣于CoCoA。统一视角还导出共识ADMM的自然原对偶间隙停止准则,并给出ADMM类方法的统一$O(1/T)$遍历收敛分析。合成回归与真实SVM数据集实验验证了预测关系,阐明调参作用,并显示调优后的ADMM可在岭正则化设定下超越CoCoA。
原文摘要 · Abstract (English)
Distributed empirical risk minimization (ERM) is often studied through two influential yet seemingly separate families of methods: CoCoA-type algorithms, derived from distributed dual coordinate ascent, and ADMM-type algorithms, derived from consensus and proximal splitting. In this paper, we investigate the connection of the two types of algorithms from a unified primal-dual perspective. We show that consensus ADMM, linearized consensus ADMM, two distributed proximal ADMM variants, and ridge-regularized CoCoA can all be written in a common update form involving a global primal variable and block dual variables. This reformulation makes several previously hidden connections explicit: For ridge-regularized ERM, CoCoA coincides with a particular proximal ADMM scheme at the level of the dual update. Moreover, consensus ADMM on the primal problem is equivalent to proximal ADMM on the dual problem under an explicit parameter mapping together with a sign reversal of the saddle objective; similar correspondences also hold for the linearized variants. These results indicates that the ADMM-type algorithms, when fine tuned, performs at least as good as CoCoA, under ridge regularized ERM problems. The unified view also yields a natural primal-dual gap stopping criterion for consensus ADMM and a unified $O(1/T)$ ergodic convergence analysis for the ADMM-type methods. Experiments on synthetic regression problems and real SVM datasets support the predicted relationships, clarify the role of tuning parameters, and show that suitably tuned ADMM variants can outperform CoCoA in the ridge-regularized setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。