统一中心化与联邦学习,用中间监督提升模型鲁棒性
OmniISR: A Unified Framework for Centralized and Federated Learning via Intermediate Supervision and Regularization

- 通过中间层监督与正则化,融合中心化与联邦学习模式
- 减少联邦学习客户端漂移,提升混合训练收敛速度22.60%
- 适用于跨区域部署的智能系统,尤其适合数据受限场景
边缘智能在全球部署面临异构法律环境:部分地区允许云端数据聚合的中心化学习(CL),而另一些地区则强制要求数据本地化的联邦学习(FL)。二者存在不兼容优化机制——中心化学习虽有无偏梯度但存在内部协变量偏移,联邦学习则面临有偏、易漂移的本地更新。为填补此空白,我们提出OmniISR,一种通过在多层隐藏层引入中间监督与正则化(ISR)信号,统一纯CL、纯FL及混合CL-FL训练模式的框架。具体包括:(i) 使用互信息(MI)作为中间监督,对齐中心化学习中的协变量偏移与联邦学习中的客户端表征漂移;(ii) 采用负熵(NE)作为中间正则项,惩罚过度自信预测,保留表征不确定性,防止设备特异性坍缩。理论上,我们推导出:(i) 统一的、无需依赖ISR的非渐近O(1/√T)收敛界,证明ISR不影响标准SGD收敛;(ii) 联邦漂移界,量化了ISR对客户端漂移的抑制作用;(iii) 梯度对齐保证,在弱偏差下确保CL与FL更新不冲突;(iv) 显式逃逸时间界,表明混合训练扩大有效随机性,加速逃离严格鞍点。大量实验表明,OmniISR在中心化与联邦范式中均显著提升性能,使CL-FL差距缩小22.60%,并在多个联邦算法上取得37/48的配对指标优势。
原文摘要 · Abstract (English)
The global deployment of edge intelligence operates across heterogeneous legal frameworks. While some regions permit centralized learning (CL) via cloud data aggregation, others enforce strict data localization, necessitating federated learning (FL). This operational dichotomy introduces two incompatible optimization regimes (i.e., unbiased global gradients yet coupled with internal covariate shift in CL versus biased, drift-prone local updates in FL), resulting in that any naive integration of the two lacks rigorous theoretical guarantees. To fill this gap, we propose OmniISR, a unified framework that fuses pure CL, pure FL, and hybrid CL-FL training modes via equipping intermediate supervision and regularization (ISR) signals at multiple hidden layers. Specifically, we propose (i) to use mutual-information (MI) as intermediate supervision to align shifting internal covariate in CL and client-drifting representations in FL, and (ii) to adopt negative-entropy (NE) as intermediate regularizer to penalize overconfident prediction, preserve representational uncertainty, and avoid device-specific collapse. On the theory side, we derive (i) a unified, ISR-agnostic, and non-asymptotic O(1/sqrt(T)) convergence bound that shows the introduced ISR does not violate standard SGD convergence, (ii) a federated drift-bound that quantifies the ISR-reduced client drift, (iii) a gradient-alignment guarantee that ensures non-conflicting CL and FL updates under mild bias, and (iv) an explicit escape-time bound that indicates that CL-FL hybrid mixing enlarges effective stochasticity and accelerates escape from strict saddles. Extensive experiments demonstrate that OmniISR consistently improves model performance in both centralized and federated paradigms, reduces the CL-FL gap by 22.60%, and yields 37/48 paired metric wins across multiple FL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。