用输出相似性替代参数平均,提升去中心化学习泛化能力
Deep-Relative-Trust-Based Diffusion for Decentralized Deep Learning
- 以神经网络输出相似度替代参数平均,构建新型去中心化学习机制
- 在稀疏拓扑下图像分类任务中显著提升模型泛化性能
- 适合资源受限、通信不畅的分布式场景,如边缘计算
去中心化学习使多个智能体无需中央聚合即可从本地数据高效学习。现有方法通常依赖参数空间的平均机制来促进一致性。我们指出,在过参数化的深度神经网络中,鼓励网络输出的一致性比参数一致性更为合适。为此,提出基于深度相对信任(DRT)的新算法DRT diffusion。该方法利用最近提出的神经网络相似性度量——深度相对信任,实现更优的去中心化学习。本文提供了该策略的收敛性分析,并通过数值实验验证其在稀疏拓扑下的图像分类任务中显著提升泛化能力。
原文摘要 · Abstract (English)
Decentralized learning strategies allow a collection of agents to learn efficiently from local data sets without the need for central aggregation or orchestration. Current decentralized learning paradigms typically rely on an averaging mechanism to encourage agreement in the parameter space. We argue that in the context of deep neural networks, which are often over-parameterized, encouraging consensus of the neural network outputs, as opposed to their parameters can be more appropriate. This motivates the development of a new decentralized learning algorithm, termed DRT diffusion, based on deep relative trust (DRT), a recently introduced similarity measure for neural networks. We provide convergence analysis for the proposed strategy, and numerically establish its benefit to generalization, especially with sparse topologies, in an image classification task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。