arXiv:2511.14456cs.DCcs.AI2025-11中稿 · publication in 3rd…被引 2

研究跨组织联邦学习中参与者失败对模型质量的影响。

Analyzing the Impact of Participant Failures in Cross-Silo Federated Learning

  • 分析失败时机与数据分布对模型的影响机制。
  • 高数据偏斜下评估结果会严重乐观,掩盖真实性能下降。
  • 适合关注联邦学习系统可靠性的研究人员与架构师。

联邦学习(FL)是一种无需共享数据即可训练机器学习(ML)模型的新范式。在跨组织(cross-silo)场景中,多个机构协作时,确保系统可靠性至关重要;但参与者可能因通信问题或配置错误等原因失败。现有研究多聚焦于资源受限的跨设备(cross-device)FL,而对跨组织场景中的失败影响研究较少。本文针对参与方较少的跨组织联邦学习,系统分析了参与者失败对模型质量的影响。重点考察失败发生时机、数据分布及评估结果的影响,发现高数据偏斜下评估结果显著乐观,掩盖真实性能损失;同时失败时机显著影响最终模型质量。研究结果为构建稳健的联邦学习系统提供了关键指导。

原文摘要 · Abstract (English)

Federated learning (FL) is a new paradigm for training machine learning (ML) models without sharing data. While applying FL in cross-silo scenarios, where organizations collaborate, it is necessary that the FL system is reliable; however, participants can fail due to various reasons (e.g., communication issues or misconfigurations). In order to provide a reliable system, it is necessary to analyze the impact of participant failures. While this problem received attention in cross-device FL where mobile devices with limited resources participate, there is comparatively little research in cross-silo FL. Therefore, we conduct an extensive study for analyzing the impact of participant failures on the model quality in the context of inter-organizational cross-silo FL with few participants. In our study, we focus on analyzing generally influential factors such as the impact of the timing and the data as well as the impact on the evaluation, which is important for deciding, if the model should be deployed. We show that under high skews the evaluation is optimistic and hides the real impact. Furthermore, we demonstrate that the timing impacts the quality of the trained model. Our results offer insights for researchers and software architects aiming to build robust FL systems.

联邦学习系统可靠性模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。