用空中叠加共享模型输出,大幅降低通信开销。
Optimal Transceiver Design in Over-the-Air Federated Distillation
- 通过空中聚合模型输出替代参数传输
- 理论推导收敛速率并求解最优功率与波束成形
- 适合边缘设备协同训练的低通信场景
人工智能的快速发展催生了联邦学习(FL),使无线设备可仅共享本地模型参数而无需传输完整数据集。然而,大模型带来的通信开销使得现有方法效率低下。本文提出一种新型空中联邦蒸馏(Over-the-Air Federated Distillation, FD)框架,融合联邦学习与知识蒸馏优势,避免传输大型模型参数。各设备不共享参数,而是通过多址信道的叠加特性,将模型输出(即知识)直接在空中聚合。研究重点在于设计收发器,以最大化学习收敛速度并满足设备功率约束。核心挑战在于学习性能分析不可行、优化问题非凸且覆盖整个训练周期。为此,本文首次推导出空中FD的收敛速率表达式,并在给定接收端组合策略下,获得发射功率与聚合估计器的闭式最优解。进一步通过半定松弛法高效求解最优接收波束成形向量,并证明该松弛无最优性损失。数值结果表明,所提方法显著降低通信开销,测试准确率仅略有下降,优于传统联邦学习基准。
原文摘要 · Abstract (English)
The rapid proliferation and growth of artificial intelligence (AI) has led to the development of federated learning (FL). FL allows wireless devices (WDs) to cooperatively learn by sharing only local model parameters, without needing to share the entire dataset. However, the emergence of large AI models has made existing FL approaches inefficient, due to the significant communication overhead required. In this paper, we propose a novel over-the-air federated distillation (FD) framework by synergizing the strength of FL and knowledge distillation to avoid the heavy local model transmission. Instead of sharing the model parameters, only the WDs' model outputs, referred to as knowledge, are shared and aggregated over-the-air by exploiting the superposition property of the multiple-access channel. We shall study the transceiver design in over-the-air FD, aiming to maximize the learning convergence rate while meeting the power constraints of the transceivers. The main challenge lies in the intractability of the learning performance analysis, as well as the non-convex nature and the optimization spanning the whole FD training period. To tackle this problem, we first derive an analytical expression of the convergence rate in over-the-air FD. Then, the closed-form optimal solutions of the WDs' transmit power and the estimator for over-the-air aggregation are obtained given the receiver combining strategy. Accordingly, we put forth an efficient approach to find the optimal receiver beamforming vector via semidefinite relaxation. We further prove that there is no optimality gap between the original and relaxed problem for the receiver beamforming design. Numerical results will show that the proposed over-the-air FD approach achieves a significant reduction in communication overhead, with only a minor compromise in testing accuracy compared to conventional FL benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。