arXiv:2507.22488cs.LGcs.AI2025-07

解决垂直联邦学习中极端数据偏移下的模型偏差问题

Proto-EVFL: Enhanced Vertical Federated Learning via Dual Prototype with Extremely Unaligned Data

  • 引入双原型机制学习类间关系,提升跨方特征对齐
  • 动态选择未对齐样本,零样本场景下性能领先基线6.97%以上
  • 适合处理多方数据严重不均衡的工业级联邦学习场景

在垂直联邦学习中,多个企业通过利用大量本地未对齐样本协同学习,以缓解对齐样本稀缺问题。然而,不同参与方间的未对齐样本可能存在极端类别不平衡,导致特征表示不足和模型预测空间受限。具体而言,类别不平衡包括参与方内不平衡与参与方间不平衡,分别引发局部模型偏差和特征贡献不一致问题。为此,我们提出Proto-EVFL,一种基于双原型的增强型垂直联邦学习框架。首先,为每方引入类别原型,学习潜在空间中的类别关系,使主动方可预测未见类别。进一步设计概率双原型学习方案,通过条件最优传输成本结合类别先验概率,动态选择未对齐样本。同时,混合先验引导模块融合局部与全局类别先验,指导选择过程。最后,采用自适应门控特征聚合策略,动态加权并聚合各参与方的本地特征,缓解特征贡献不一致问题。我们证明Proto-EVFL是首个在垂直联邦学习中实现双层优化的框架,具有1/√T的收敛速率。在多种数据集上的实验验证了其优越性,即使在存在一个未见类别的零样本场景下,性能仍优于基线至少6.97%。

原文摘要 · Abstract (English)

In vertical federated learning (VFL), multiple enterprises address aligned sample scarcity by leveraging massive locally unaligned samples to facilitate collaborative learning. However, unaligned samples across different parties in VFL can be extremely class-imbalanced, leading to insufficient feature representation and limited model prediction space. Specifically, class-imbalanced problems consist of intra-party class imbalance and inter-party class imbalance, which can further cause local model bias and feature contribution inconsistency issues, respectively. To address the above challenges, we propose Proto-EVFL, an enhanced VFL framework via dual prototypes. We first introduce class prototypes for each party to learn relationships between classes in the latent space, allowing the active party to predict unseen classes. We further design a probabilistic dual prototype learning scheme to dynamically select unaligned samples by conditional optimal transport cost with class prior probability. Moreover, a mixed prior guided module guides this selection process by combining local and global class prior probabilities. Finally, we adopt an \textit{adaptive gated feature aggregation strategy} to mitigate feature contribution inconsistency by dynamically weighting and aggregating local features across different parties. We proved that Proto-EVFL, as the first bi-level optimization framework in VFL, has a convergence rate of 1/\sqrt T. Extensive experiments on various datasets validate the superiority of our Proto-EVFL. Even in a zero-shot scenario with one unseen class, it outperforms baselines by at least 6.97%

联邦学习数据偏移原型学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。