arXiv:2411.00388cs.GTcs.LG2024-11被引 2

改进数据价值评估方法,让数据贡献更真实反映其实际作用

Towards Data Valuation via Asymmetric Data Shapley

  • 打破传统方法对数据同质的假设,引入数据结构差异
  • 提出基于k近邻的高效算法,实现精确计算
  • 适用于真实场景中的数据交易与模型优化

随着数据成为技术与经济进步的关键驱动力,如何在算法决策中准确量化其价值成为核心挑战。传统的数据Shapley值虽被广泛用于评估监督学习中单个数据源的贡献,但其对称性假设忽略了现实数据集中的复杂结构与依赖关系。为此,本文提出非对称数据Shapley框架,可灵活融入数据内在结构,实现结构感知的数据价值评估。同时设计了一种基于k近邻的高效算法,实现精确计算。实验验证了该框架在多种机器学习任务和数据市场场景中的实用性。代码已公开于https://github.com/xzheng01/Asymmetric-Data-Shapley。

原文摘要 · Abstract (English)

As data emerges as a vital driver of technological and economic advancements, a key challenge is accurately quantifying its value in algorithmic decision-making. The Shapley value, a well-established concept from cooperative game theory, has been widely adopted to assess the contribution of individual data sources in supervised machine learning. However, its symmetry axiom assumes all players in the cooperative game are homogeneous, which overlooks the complex structures and dependencies present in real-world datasets. To address this limitation, we extend the traditional data Shapley framework to asymmetric data Shapley, making it flexible enough to incorporate inherent structures within the datasets for structure-aware data valuation. We also introduce an efficient $k$-nearest neighbor-based algorithm for its exact computation. We demonstrate the practical applicability of our framework across various machine learning tasks and data market contexts. The code is available at: https://github.com/xzheng01/Asymmetric-Data-Shapley.

数据估值Shapley值机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。