arXiv:2412.17317cs.LGcs.SE2024-12被引 7

在保护隐私前提下,提升跨项目缺陷预测的准确性。

Better Knowledge Enhancement for Privacy-Preserving Cross-Project Defect Prediction

  • 通过联邦学习与知识蒸馏融合,缓解数据异构问题。
  • 在19个开源项目上,性能显著优于现有基线方法。
  • 适合关注数据隐私与模型泛化能力的研究者。

跨项目缺陷预测(CPDP)在利用其他项目数据构建可靠缺陷预测模型时面临挑战,尤其当数据所有者关注隐私安全时。近年来,联邦学习(FL)成为一种新兴范式,可在不共享原始数据的前提下实现多方协作训练全局模型,保障隐私。然而,不同公司或组织间专有项目带来的数据异构性会干扰模型训练。本文研究在联邦学习框架下,针对数据异构性的隐私保护跨项目缺陷预测问题。为此,提出一种名为FedDP的新颖知识增强方法,包含两个简单但有效的解决方案:局部异构感知与全局知识蒸馏。具体地,使用开源项目数据作为蒸馏数据集,通过异构感知的本地模型集成优化全局模型。在两个数据集共19个项目的实验中,所提方法显著优于基线方法。

原文摘要 · Abstract (English)

Cross-Project Defect Prediction (CPDP) poses a non-trivial challenge to construct a reliable defect predictor by leveraging data from other projects, particularly when data owners are concerned about data privacy. In recent years, Federated Learning (FL) has become an emerging paradigm to guarantee privacy information by collaborative training a global model among multiple parties without sharing raw data. While the direct application of FL to the CPDP task offers a promising solution to address privacy concerns, the data heterogeneity arising from proprietary projects across different companies or organizations will bring troubles for model training. In this paper, we study the privacy-preserving cross-project defect prediction with data heterogeneity under the federated learning framework. To address this problem, we propose a novel knowledge enhancement approach named FedDP with two simple but effective solutions: 1. Local Heterogeneity Awareness and 2. Global Knowledge Distillation. Specifically, we employ open-source project data as the distillation dataset and optimize the global model with the heterogeneity-aware local model ensemble via knowledge distillation. Experimental results on 19 projects from two datasets demonstrate that our method significantly outperforms baselines.

缺陷预测联邦学习知识蒸馏隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。