解析大规模神经网络在去中心化训练中的学习动态
Peer-to-Peer Learning Dynamics of Wide Neural Networks
- 基于神经正切核理论分析去中心化梯度下降的训练过程
- 准确预测了分类任务中参数与误差的演化轨迹
- 适合研究分布式机器学习与边缘计算的开发者
去中心化学习是新兴的分布式边缘设备协同训练深度神经网络的框架,可在无中央服务器的情况下实现隐私保护。针对智能城市等场景,神经网络架构与超参数设计难以在实际部署中调优,亟需对非凸神经网络在去中心化环境中的训练动态进行刻画。本文首次对使用主流分布式梯度下降(DGD)算法训练的宽神经网络学习动态提供显式表征。结果结合了神经正切核(NTK)理论与分布式学习共识的前期成果,并通过大量实验验证了对分类任务中参数和误差动态的精确预测能力。
原文摘要 · Abstract (English)
Peer-to-peer learning is an increasingly popular framework that enables beyond-5G distributed edge devices to collaboratively train deep neural networks in a privacy-preserving manner without the aid of a central server. Neural network training algorithms for emerging environments, e.g., smart cities, have many design considerations that are difficult to tune in deployment settings -- such as neural network architectures and hyperparameters. This presents a critical need for characterizing the training dynamics of distributed optimization algorithms used to train highly nonconvex neural networks in peer-to-peer learning environments. In this work, we provide an explicit characterization of the learning dynamics of wide neural networks trained using popular distributed gradient descent (DGD) algorithms. Our results leverage both recent advancements in neural tangent kernel (NTK) theory and extensive previous work on distributed learning and consensus. We validate our analytical results by accurately predicting the parameter and error dynamics of wide neural networks trained for classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。