通过筛选可信的补全样本,提升垂直联邦学习在少量重叠数据下的性能。
Reliable Imputed-Sample Assisted Vertical Federated Learning
- 用证据理论评估补全样本的不确定性,仅保留低不确定性的样本参与训练。
- 在仅有1%重叠样本时,CIFAR-10上准确率提升48%。
- 适合数据重叠少、需挖掘非重叠样本的垂直联邦学习场景。
垂直联邦学习(VFL)允许多方在不共享原始数据的前提下协同建模。现有方法主要依赖各方数据的重叠样本,但受限于重叠样本数量有限,大量非重叠样本未被利用。部分先前工作尝试对缺失值进行补全,但常忽略补全样本的质量。为此,我们提出可靠的补全样本辅助(RISA)VFL框架,通过筛选高质量补全样本以有效利用非重叠样本。具体地,在补全非重叠样本后,引入证据理论估计其不确定性,并仅选取低不确定性样本参与训练。实验表明,RISA在两个常用数据集上均取得显著性能提升,尤其在重叠样本极少时表现突出:在仅1%重叠样本的CIFAR-10上,准确率提升48%。
原文摘要 · Abstract (English)
Vertical Federated Learning (VFL) is a well-known FL variant that enables multiple parties to collaboratively train a model without sharing their raw data. Existing VFL approaches focus on overlapping samples among different parties, while their performance is constrained by the limited number of these samples, leaving numerous non-overlapping samples unexplored. Some previous work has explored techniques for imputing missing values in samples, but often without adequate attention to the quality of the imputed samples. To address this issue, we propose a Reliable Imputed-Sample Assisted (RISA) VFL framework to effectively exploit non-overlapping samples by selecting reliable imputed samples for training VFL models. Specifically, after imputing non-overlapping samples, we introduce evidence theory to estimate the uncertainty of imputed samples, and only samples with low uncertainty are selected. In this way, high-quality non-overlapping samples are utilized to improve VFL model. Experiments on two widely used datasets demonstrate the significant performance gains achieved by the RISA, especially with the limited overlapping samples, e.g., a 48% accuracy gain on CIFAR-10 with only 1% overlapping samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。