arXiv:2504.19882cs.CV2025-04中稿 · ed被引 5

用因果方法生成更丰富的数据,提升联邦学习在分布外样本上的泛化能力。

Federated Out-of-Distribution Generalization: A Causal Augmentation View

  • 通过因果区域定位分离图像背景与主体,解耦伪相关性。
  • 基于因果特征生成反事实样本,显著提升数据多样性与模型鲁棒性。
  • 无需客户端间数据共享,保护隐私,适合医疗等敏感领域应用。

联邦学习旨在整合多源信息协同建模,使模型能跨客户端数据泛化。现有方法常通过知识蒸馏或数据增强缓解客户端间数据偏差的负面影响,但教师模型在分布外样本上表现有限,且增强数据与原始数据存在质量差距,难以利用丰富上下文信息。为此,本文提出联邦因果增强方法 FedCAug,通过因果启发式数据增强打破属性与类别间的虚假关联。具体地,设计因果区域定位模块精准识别并分离图像中的背景与物体,为因果增强提供丰富上下文;同时设计因果启发式数据增强模块,融合因果特征与客户端内上下文生成反事实样本。该过程不需客户端间信息交换,有效保护数据隐私。在三个数据集上的实验表明,FedCAug显著降低模型对背景的依赖,性能优于现有先进方法。

原文摘要 · Abstract (English)

Federated learning aims to collaboratively model by integrating multi-source information to obtain a model that can generalize across all client data. Existing methods often leverage knowledge distillation or data augmentation to mitigate the negative impact of data bias across clients. However, the limited performance of teacher models on out-of-distribution samples and the inherent quality gap between augmented and original data hinder their effectiveness and they typically fail to leverage the advantages of incorporating rich contextual information. To address these limitations, this paper proposes a Federated Causal Augmentation method, termed FedCAug, which employs causality-inspired data augmentation to break the spurious correlation between attributes and categories. Specifically, it designs a causal region localization module to accurately identify and decouple the background and objects in the image, providing rich contextual information for causal data augmentation. Additionally, it designs a causality-inspired data augmentation module that integrates causal features and within-client context to generate counterfactual samples. This significantly enhances data diversity, and the entire process does not require any information sharing between clients, thereby contributing to the protection of data privacy. Extensive experiments conducted on three datasets reveal that FedCAug markedly reduces the model's reliance on background to predict sample labels, achieving superior performance compared to state-of-the-art methods.

联邦学习因果推理数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。