arXiv:2505.11035cs.LG2025-05中稿 · ICLR

提出首个统一处理多机构数据对齐、标签缺失的垂直联邦学习框架

Deep Latent Variable Model based Vertical Federated Learning with Flexible Alignment and Labeling Scenarios

  • 将数据对齐差异视为缺失数据问题,构建统一建模框架
  • 168组实验中160次优于基线,平均性能领先9.6个百分点
  • 适合跨机构合作且数据不全、对齐不一致的复杂场景

联邦学习(FL)通过不共享原始数据实现协作学习,备受关注。其中,垂直联邦学习(VFL)针对多个机构持有的用户特征数据进行协作,每方掌握互补信息。然而现有方法常受限于参与方数量少、数据完全对齐或仅使用标注数据等假设。本文将VFL中的对齐缺口重新定义为缺失数据问题,提出一个统一框架,可在任意对齐与标注场景下支持训练与推理,并兼容多种缺失机制。在涵盖四个基准数据集、六种训练期缺失模式、七种测试期缺失模式的168组配置实验中,该方法在160次中超越所有基线,平均性能领先第二名9.6个百分点。据我们所知,这是首个能同时处理任意数据对齐、无标签数据与多方协作的VFL框架。

原文摘要 · Abstract (English)

Federated learning (FL) has attracted significant attention for enabling collaborative learning without exposing private data. Among the primary variants of FL, vertical federated learning (VFL) addresses feature-partitioned data held by multiple institutions, each holding complementary information for the same set of users. However, existing VFL methods often impose restrictive assumptions such as a small number of participating parties, fully aligned data, or only using labeled data. In this work, we reinterpret alignment gaps in VFL as missing data problems and propose a unified framework that accommodates both training and inference under arbitrary alignment and labeling scenarios, while supporting diverse missingness mechanisms. In the experiments on 168 configurations spanning four benchmark datasets, six training-time missingness patterns, and seven testing-time missingness patterns, our method outperforms all baselines in 160 cases with an average gap of 9.6 percentage points over the next-best competitors. To the best of our knowledge, this is the first VFL framework to jointly handle arbitrary data alignment, unlabeled data, and multi-party collaboration all at once.

联邦学习垂直联邦数据对齐缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。