构建首个真实遥感联邦学习数据集与基准,支持大规模、异构场景评估。
FedRS-Bench: Realistic Federated Learning Datasets and Benchmarks in Remote Sensing
- 基于8个真实遥感数据集构建135个客户端,模拟真实分布式数据分布。
- 实验显示联邦学习在异构条件下仍能提升模型性能,揭示不同方法的优劣。
- 适合关注遥感联邦学习、跨域数据协作的研究者使用。
遥感(RS)图像以空前规模生成,但地理与机构分布分散,受数据共享限制和隐私担忧影响,集中式模型训练困难。联邦学习(FL)通过在去中心化数据源间协作训练模型而不暴露原始数据,提供了解决方案。然而,当前缺乏真实的遥感联邦学习数据集与基准。以往研究多依赖手动划分单一数据集,无法捕捉真实遥感数据的异构性与规模,且实验设置不统一,阻碍公平比较。为此,我们提出真实联邦遥感数据集 FedRS,包含8个覆盖多种传感器与分辨率的数据集,构建135个客户端,反映实际操作场景。每个客户端数据来自同一来源,具备标签分布偏斜、数据量不均衡、跨客户端域异构等真实联邦特性,体现实际挑战并支持大规模联邦学习方法评估。基于 FedRS,我们实现10种基线联邦学习算法与评估指标,构建全面的 FedRS-Bench。实验表明,联邦学习在独立数据孤岛之外能持续提升模型性能,同时揭示不同方法在客户端异构性与可用性变化下的性能权衡。我们希望 FedRS-Bench 能推动遥感领域大规模、真实联邦学习研究,提供标准化测试平台,促进未来工作的公平比较。代码与数据集已开源:https://fedrs-bench.github.io/
原文摘要 · Abstract (English)
Remote sensing (RS) images are usually produced at an unprecedented scale, yet they are geographically and institutionally distributed, making centralized model training challenging due to data-sharing restrictions and privacy concerns. Federated learning (FL) offers a solution by enabling collaborative model training across decentralized RS data sources without exposing raw data. However, there lacks a realistic federated dataset and benchmark in RS. Prior works typically rely on manually partitioned single dataset, which fail to capture the heterogeneity and scale of real-world RS data, and often use inconsistent experimental setups, hindering fair comparison. To address this gap, we propose a realistic federated RS dataset, termed FedRS. FedRS consists of eight datasets that cover various sensors and resolutions and builds 135 clients, which is representative of realistic operational scenarios. Data for each client come from the same source, exhibiting authentic federated properties such as skewed label distributions, imbalanced client data volumes, and domain heterogeneity across clients. These characteristics reflect practical challenges in federated RS and support evaluation of FL methods at scale. Based on FedRS, we implement 10 baseline FL algorithms and evaluation metrics to construct the comprehensive FedRS-Bench. The experimental results demonstrate that FL can consistently improve model performance over training on isolated data silos, while revealing performance trade-offs of different methods under varying client heterogeneity and availability conditions. We hope FedRS-Bench will accelerate research on large-scale, realistic FL in RS by providing a standardized, rich testbed and facilitating fair comparisons across future works. The source codes and dataset are available at https://fedrs-bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。