一种可保护隐私且抗数据增强的数据估值方法
Private, Augmentation-Robust and Task-Agnostic Data Valuation Approach for Data Marketplace
- 通过计算买家与卖家数据分布距离评估数据价值
- 无需获取完整数据即可完成估值,且结果对数据增强不变
- 适合注重隐私与数据质量的平台方和数据买家
在数据交易市场中,买家需评估待购数据的价值,这是关键挑战。本文提出一种名为PriArTa的任务无关数据估值方法,通过计算买家现有数据与卖家数据之间的分布距离,判断新数据对自身数据集的增益效果。该方法通信高效,买家无需获取卖家全部数据,只需请求卖家对数据进行特定预处理并返回结果,再结合评分机制完成估值。预处理设计确保买家可计算得分的同时保护卖家数据隐私,降低交易前信息泄露风险。其核心优势在于对常见数据变换具有鲁棒性,保证估值一致性,避免购买冗余数据。实验在真实图像数据集上验证了其有效性,证明该方法可在保护隐私的前提下实现抗数据增强的数据估值。
原文摘要 · Abstract (English)
Evaluating datasets in data marketplaces, where the buyer aim to purchase valuable data, is a critical challenge. In this paper, we introduce an innovative task-agnostic data valuation method called PriArTa which is an approach for computing the distance between the distribution of the buyer's existing dataset and the seller's dataset, allowing the buyer to determine how effectively the new data can enhance its dataset. PriArTa is communication-efficient, enabling the buyer to evaluate datasets without needing access to the entire dataset from each seller. Instead, the buyer requests that sellers perform specific preprocessing on their data and then send back the results. Using this information and a scoring metric, the buyer can evaluate the dataset. The preprocessing is designed to allow the buyer to compute the score while preserving the privacy of each seller's dataset, mitigating the risk of information leakage before the purchase. A key feature of PriArTa is its robustness to common data transformations, ensuring consistent value assessment and reducing the risk of purchasing redundant data. The effectiveness of PriArTa is demonstrated through experiments on real-world image datasets, showing its ability to perform privacy-preserving, augmentation-robust data valuation in data marketplaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。