用博弈论方法评估数据掩码中各特征的隐私风险与价值,实现精准平衡。
Shapley-Value-Based Feature Attribution for Data Masking
- 基于谢林值计算特征对隐私泄露和数据效用的贡献度。
- 在不依赖特定模型或指标的前提下,降低泄露风险并保持数据可用性。
- 适用于各类数据掩码技术,适合隐私保护研究者使用。
尽管个人数据广泛共享带来诸多好处,但也引发消费者、企业和政策制定者严重的隐私担忧。本文提出一种新颖框架,将基于谢林值的特征归因方法应用于数据隐私领域,捕捉数据隐私的两个关键维度——披露风险与数据效用。该框架从整体视角出发,通过公平的谢林值特征归因方法实现数据掩码的综合评估。与现有文献多关注数据集层面的风险-效用权衡不同,本框架在特征层面解决该问题。此外,该框架对数据掩码方法、统计与机器学习方法以及效用与风险评估指标均无偏好。实验结果表明,所提方法能有效降低披露风险,同时保持数据效用。
原文摘要 · Abstract (English)
Despite its many benefits, widespread access to individuals' personal data also causes severe privacy concerns for consumers, companies, and policymakers. This study proposes a novel framework that adapts the Shapley-value-based feature attribution approach to the problem domain of data privacy by capturing the two crucial dimensions of data privacy---disclosure risk and data utility. Our proposed framework takes a holistic view of data masking through a fair feature attribution approach based on Shapley values. Different from the existing literature that mostly focuses on the risk-utility tradeoff at the dataset level, the proposed framework addresses the tradeoff at the feature level. Furthermore, the proposed framework is agnostic to data masking methods, statistical and machine learning methods, and data utility and disclosure risk evaluation metrics. Experimental results show that our proposed method can effectively reduce disclosure risk while preserving data utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。