用机器遗忘估算数据价值,高效又隐私安全。
Losing is for Cherishing: Data Valuation Based on Machine Unlearning and Shapley Value
- 基于机器遗忘和蒙特卡洛采样,避免重训练
- 在多个数据集上准确率接近顶尖方法,计算量降低数量级
- 适合大模型数据市场,支持部分数据估值
大规模模型的兴起加剧了对高效数据估值方法的需求,以量化单个数据提供者的贡献。传统方法如基于博弈论的谢尔普利值和基于影响函数的技术,面临高昂计算成本或需访问完整数据与模型训练细节,难以实现部分数据估值。为此,我们提出Unlearning Shapley框架,利用机器遗忘技术高效估算数据价值。通过从预训练模型中移除目标数据,并测量在可访问测试集上的性能变化,结合蒙特卡洛采样计算谢尔普利值,无需重训练且不依赖完整数据。关键优势在于支持全量与部分数据估值,适用于大型模型(如LLMs)和数据市场。在基准数据集和大规模文本语料上的实验表明,该方法在保持与先进方法相当准确性的同时,计算开销降低数量级。进一步分析证实估计值与数据子集真实影响高度相关,验证其在现实场景中的可靠性。本工作弥合了数据估值理论与实际部署之间的差距,为现代AI生态系统提供了一种可扩展、符合隐私要求的解决方案。
原文摘要 · Abstract (English)
The proliferation of large models has intensified the need for efficient data valuation methods to quantify the contribution of individual data providers. Traditional approaches, such as game-theory-based Shapley value and influence-function-based techniques, face prohibitive computational costs or require access to full data and model training details, making them hardly achieve partial data valuation. To address this, we propose Unlearning Shapley, a novel framework that leverages machine unlearning to estimate data values efficiently. By unlearning target data from a pretrained model and measuring performance shifts on a reachable test set, our method computes Shapley values via Monte Carlo sampling, avoiding retraining and eliminating dependence on full data. Crucially, Unlearning Shapley supports both full and partial data valuation, making it scalable for large models (e.g., LLMs) and practical for data markets. Experiments on benchmark datasets and large-scale text corpora demonstrate that our approach matches the accuracy of state-of-the-art methods while reducing computational overhead by orders of magnitude. Further analysis confirms a strong correlation between estimated values and the true impact of data subsets, validating its reliability in real-world scenarios. This work bridges the gap between data valuation theory and practical deployment, offering a scalable, privacy-compliant solution for modern AI ecosystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。