提出异步对等网络中知识蒸馏的收敛理论,解决模型差异下的联合训练难题。
Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network

- 将共识从参数空间转移到输出空间,用软标签实现跨模型协作
- 函数偏差与驻留性以1/(ηT)速率收敛至η+ B_f²+ ζ_f²邻域
- 实验证明蒸馏使函数差异缩小40-61倍,适合异构设备协同训练
去中心化、无服务器学习正连接具有不同架构的设备,但标准的去中心化SGD因参数量不一致无法平均而失效。知识蒸馏(KD)通过交换软预测而非权重绕过此问题,然而完全去中心化、异步对等(P2P)KD的收敛理论尚缺。本文提供理论:将共识从参数空间转移到函数(输出)空间——一个KD事件是逻辑空间中预测分布的几何收缩算子,我们在参考测度上的预测希尔伯特空间中分析其性质。在标准光滑性/方差假设及两个可实现性假设下(一个连接参数SGD与函数步,一个控制受限任务与蒸馏对齐),时间平均的函数驻留性与函数空间分歧以 $O(1/( heta T))$ 速率收敛至 $O( heta)+O(B_f^2)+O( heta_f^2)$ 邻域。其中 $B_f$ 表示任务最优到同伴可达类的距离,$ heta_f$ 衡量持续的局部任务异质性。在同质、宽度异构及混合家族网络实验中,KD使函数分歧降低40-61倍,孤立训练则无改善。采样驻留性诊断显示主实验后期瞬态指数为0.99-1.90,四点步长扫描验证了预测的邻域权衡现象。
原文摘要 · Abstract (English)
Decentralized, serverless learning increasingly connects devices running different architectures, where the standard tool, decentralized SGD, is undefined as models with different parameter counts cannot be averaged. Knowledge distillation (KD) exchanges soft predictions rather than weights and sidesteps this obstacle, yet convergence theory for fully decentralized, asynchronous peer-to-peer (P2P) KD is lacking. We provide one, relocating consensus from parameter space to function (output) space: a KD event is a geometric contraction operator in logit space on the peers' predictive distributions, which we analyse in the Hilbert space of predictions on a reference measure. Under standard smoothness/variance assumptions and two realizability assumptions, one bridging parameter SGD to the functional step and one controlling restricted task/KD alignment, the time-averaged functional stationarity and function-space disagreement converge at rate $O(1/(\eta T))$ to an $O(\eta)+O(B_f^2)+O(\zeta_f^2)$ neighbourhood. Here $B_f$ is the distance from the task optimum to the peers' reachable classes and $\zeta_f$ measures persistent local-task heterogeneity. Across homogeneous, width-heterogeneous, and mixed-family networks of the experiments, KD contracts function disagreement by $40-61\times$, while isolated training does not. The sampled stationarity diagnostic has late transient exponents $0.99-1.90$ on the shared-skeleton main runs, and the four-point step-size sweep exhibits the predicted transient: neighbourhood tradeoff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。