arXiv:2409.06123cs.LG2024-09被引 5

解决表格数据孤岛的对比联邦学习,不共享数据也能提升模型性能。

Contrastive Federated Learning with Tabular Data Silos

  • 在各数据孤岛内构建对比表示,通过联邦学习聚合知识。
  • 实验显示性能优于现有方法,且保持隐私安全。
  • 适合有数据隔离需求的医疗、金融等敏感领域应用。

垂直划分的数据孤岛因数据分散、样本对齐困难和严格隐私要求,难以有效学习。联邦学习被提出作为解决方案,但跨孤岛的样本错位常阻碍模型最佳性能,且需共享数据以实现性能提升,这会破坏隐私。我们提出对比联邦学习(CFL),用于处理存在样本错位的数据孤岛,无需共享原始或代表性数据即可维护隐私。CFL 在每个孤岛内本地获取数据的对比表示,并通过联邦学习算法从其他孤岛聚合知识。实验表明,CFL 解决了现有算法在数据孤岛上的局限性,在表格对比学习中表现更优,且未降低隐私保护水平。

原文摘要 · Abstract (English)

Learning from vertical partitioned data silos is challenging due to the segmented nature of data, sample misalignment, and strict privacy concerns. Federated learning has been proposed as a solution. However, sample misalignment across silos often hinders optimal model performance and suggests data sharing within the model, which breaks privacy. Our proposed solution is Contrastive Federated Learning with Tabular Data Silos (CFL), which offers a solution for data silos with sample misalignment without the need for sharing original or representative data to maintain privacy. CFL begins with local acquisition of contrastive representations of the data within each silo and aggregates knowledge from other silos through the federated learning algorithm. Our experiments demonstrate that CFL solves the limitations of existing algorithms for data silos and outperforms existing tabular contrastive learning. CFL provides performance improvements without loosening privacy.

联邦学习对比学习数据孤岛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。