新模型同时保护隐私与数据效用,提升机器学习性能。
Multi-Objective Optimization-Based Anonymization of Structured Data for Machine Learning Application
- 多目标优化兼顾隐私保护与信息损失最小化。
- 相比旧方法,受攻击风险个体减少,机器学习效果相当。
- 适用于需要高隐私保障的数据共享场景。
机构收集了大量数据,但常缺乏充分挖掘洞察的能力,因此越来越多地将数据共享给外部专家以获取价值。然而,这一做法带来显著隐私风险。尽管已有多种隐私保护技术被提出,但这些方法常导致数据效用下降,影响机器学习(ML)模型性能。本研究发现现有优化模型在处理类别变量及跨数据集评估方面存在关键局限。为此,我们提出一种新的多目标优化模型,同时最小化信息损失并最大化对攻击的防护能力。该模型在多个数据集上进行实证验证,并与两种现有算法对比。评估指标包括信息损失、易受链接攻击或同质性攻击的个体数量,以及匿名化后的机器学习性能。结果表明,所提模型在降低信息损失和有效缓解攻击风险方面优于对比算法,在某些情况下显著减少了潜在受害个体数;同时保持与原始数据或其它匿名化方法相当的机器学习表现。研究结果表明,该框架在隐私保护与数据效用之间实现了显著平衡,为数据共享提供了可扩展的解决方案。
原文摘要 · Abstract (English)
Organizations are collecting vast amounts of data, but they often lack the capabilities needed to fully extract insights. As a result, they increasingly share data with external experts, such as analysts or researchers, to gain value from it. However, this practice introduces significant privacy risks. Various techniques have been proposed to address privacy concerns in data sharing. However, these methods often degrade data utility, impacting the performance of machine learning (ML) models. Our research identifies key limitations in existing optimization models for privacy preservation, particularly in handling categorical variables, and evaluating effectiveness across diverse datasets. We propose a novel multi-objective optimization model that simultaneously minimizes information loss and maximizes protection against attacks. This model is empirically validated using diverse datasets and compared with two existing algorithms. We assess information loss, the number of individuals subject to linkage or homogeneity attacks, and ML performance after anonymization. The results indicate that our model achieves lower information loss and more effectively mitigates the risk of attacks, reducing the number of individuals susceptible to these attacks compared to alternative algorithms in some cases. Additionally, our model maintains comparable ML performance relative to the original data or data anonymized by other methods. Our findings highlight significant improvements in privacy protection and ML model performance, offering a comprehensive and extensible framework for balancing privacy and utility in data sharing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。