无需共享数据,通过交换吉布斯分布即可实现中心化学习效果。
Decentralized Machine Learning with Centralized Performance Guarantees via Gibbs Algorithms
- 客户端用带相对熵正则的优化框架,通过前后向通信共享本地吉布斯测度。
- 仅需按本地样本量缩放正则化因子,即可达到与中心化训练相同性能。
- 适合关注隐私保护、模型协作但不想传数据的研究者或系统设计者。
本文首次证明,在不共享本地数据集的前提下,可实现中心化学习性能。当客户端采用带相对熵正则的经验风险最小化(ERM-RER)框架,并建立客户端间的前向-后向通信时,只需共享本地生成的吉布斯测度,即可获得与全局集中式ERM-RER同等的性能。核心思想是:客户端k生成的吉布斯测度作为客户端k+1的参考测度,从而以一种合理方式编码先验信息。特别地,在去中心化设置中实现中心化性能,需将正则化因子按本地样本量进行特定缩放。该结果为新型去中心化学习范式开辟了道路,将协作策略从数据共享转向通过模型空间的参考测度共享局部归纳偏置。
原文摘要 · Abstract (English)
In this paper, it is shown, for the first time, that centralized performance is achievable in decentralized learning without sharing the local datasets. Specifically, when clients adopt an empirical risk minimization with relative-entropy regularization (ERM-RER) learning framework and a forward-backward communication between clients is established, it suffices to share the locally obtained Gibbs measures to achieve the same performance as that of a centralized ERM-RER with access to all the datasets. The core idea is that the Gibbs measure produced by client~$k$ is used, as reference measure, by client~$k+1$. This effectively establishes a principled way to encode prior information through a reference measure. In particular, achieving centralized performance in the decentralized setting requires a specific scaling of the regularization factors with the local sample sizes. Overall, this result opens the door to novel decentralized learning paradigms that shift the collaboration strategy from sharing data to sharing the local inductive bias via the reference measures over the set of models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。