构建联邦学习评估新基准,聚焦适应性、可信性与推理能力
ATR-Bench: A Federated Learning Benchmark for Adaptation, Trust, and Reasoning
- 从适应性、可信性、推理三维度构建统一评估框架
- 覆盖异构客户端与对抗环境下的方法与数据集评测
- 为真实场景下的联邦学习研究提供可复现的基准
联邦学习(FL)作为一种在保护数据隐私的前提下实现分布式协作训练的范式,正日益受到关注。随着其应用扩展,大量技术被提出以应对实际挑战,但缺乏对关键维度的标准化评估,阻碍了方法间的系统性比较与进展。本文提出ATR-Bench,一个涵盖适应性、可信性与推理三个核心维度的统一评估框架。我们深入分析各维度的概念基础、任务定义及开放问题,并对代表性方法和数据集进行了广泛评测,尤其在异构客户端和对抗/不可靠环境中的可信性方面。由于推理能力在联邦学习中尚无可靠度量与模型,该维度仅提供文献驱动的洞见。ATR-Bench为具有现实意义的联邦学习系统性评估奠定基础,完整代码库将公开,并维护一个持续跟踪最新研究成果的资源库。
原文摘要 · Abstract (English)
Federated Learning (FL) has emerged as a promising paradigm for collaborative model training while preserving data privacy across decentralized participants. As FL adoption grows, numerous techniques have been proposed to tackle its practical challenges. However, the lack of standardized evaluation across key dimensions hampers systematic progress and fair comparison of FL methods. In this work, we introduce ATR-Bench, a unified framework for analyzing federated learning through three foundational dimensions: Adaptation, Trust, and Reasoning. We provide an in-depth examination of the conceptual foundations, task formulations, and open research challenges associated with each theme. We have extensively benchmarked representative methods and datasets for adaptation to heterogeneous clients and trustworthiness in adversarial or unreliable environments. Due to the lack of reliable metrics and models for reasoning in FL, we only provide literature-driven insights for this dimension. ATR-Bench lays the groundwork for a systematic and holistic evaluation of federated learning with real-world relevance. We will make our complete codebase publicly accessible and a curated repository that continuously tracks new developments and research in the FL literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。