针对卫星数据非独立同分布问题,提出聚类联邦学习方法提升遥感图像识别精度。
Towards Satellite Non-IID Imagery: A Spectral Clustering-Assisted Federated Learning Approach
- 用谱聚类根据模型更新相似性动态分组卫星客户端,解决数据异质性挑战。
- 在SAT4数据集上,准确率分别比pFedSD、FedProx等方法高1.01倍至2.15倍。
- 适合低轨卫星遥感场景,尤其适用于带宽受限、数据分布不均的边缘计算任务。
低地球轨道(LEO)卫星能够获取丰富的地球观测数据(EOD),支持多种物联网应用。然而,要实现高效的EOD处理机制,亟需解决两个问题:一是由于卫星与地面站连接间歇,无法将大尺寸观测数据传回地面;二是卫星数据呈现非独立同分布(non-IID)特性。本文提出一种基于轨道的谱聚类辅助分簇联邦自知识蒸馏(OSC-FSKD)方法,应用于每条轨道的LEO卫星星座,保留联邦学习无需上传原始数据的优势。具体地,引入基于归一化拉普拉斯的谱聚类(NLSC)于联邦学习中,每轮动态依据模型更新的余弦相似性将客户端分为若干簇,以应对non-IID数据带来的挑战。同时,采用自知识蒸馏构建本地客户端,用最新更新的本地模型指导当前训练。实验表明,在SAT4数据集上,该方法的观测准确率分别较pFedSD、FedProx、FedAU、FedALA高出1.01倍、2.15倍、1.10倍和1.03倍。该方法在其他数据集上也表现出优越性。
原文摘要 · Abstract (English)
Low Earth orbit (LEO) satellites are capable of gathering abundant Earth observation data (EOD) to enable different Internet of Things (IoT) applications. However, to accomplish an effective EOD processing mechanism, it is imperative to investigate: 1) the challenge of processing the observed data without transmitting those large-size data to the ground because the connection between the satellites and the ground stations is intermittent, and 2) the challenge of processing the non-independent and identically distributed (non-IID) satellite data. In this paper, to cope with those challenges, we propose an orbit-based spectral clustering-assisted clustered federated self-knowledge distillation (OSC-FSKD) approach for each orbit of an LEO satellite constellation, which retains the advantage of FL that the observed data does not need to be sent to the ground. Specifically, we introduce normalized Laplacian-based spectral clustering (NLSC) into federated learning (FL) to create clustered FL in each round to address the challenge resulting from non-IID data. Particularly, NLSC is adopted to dynamically group clients into several clusters based on cosine similarities calculated by model updates. In addition, self-knowledge distillation is utilized to construct each local client, where the most recent updated local model is used to guide current local model training. Experiments demonstrate that the observation accuracy obtained by the proposed method is separately 1.01x, 2.15x, 1.10x, and 1.03x higher than that of pFedSD, FedProx, FedAU, and FedALA approaches using the SAT4 dataset. The proposed method also shows superiority when using other datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。