arXiv:2410.04386cs.LG2024-10NeurIPS被引 8

评估数据分布价值,让买家用小样本就能选对数据源。

Data Distribution Valuation

  • 基于MMD构建分布估值方法,可从样本比较分布价值。
  • 在多个真实数据集上表现优于现有方法,样本效率高。
  • 适合数据采购、市场定价等需要评估数据分布的场景。

数据估值是一类用于量化数据在数据市场等应用中价值的技术。现有方法仅针对离散数据集定义价值,但在许多场景中,用户更关心数据分布本身的价值。例如,买家需通过各供应商提供的少量预览样本,判断哪个数据分布更有价值并决定购买。核心问题是:如何从样本中比较不同数据分布的价值?本文在Huber异质性假设下,提出一种基于最大均值差异(MMD)的估值方法,实现了理论上合理且可操作的分布比较策略。实验表明,该方法在多个真实数据集(如网络入侵检测、信用卡欺诈检测)和下游任务(分类、回归)上,相比现有基线表现出更高的样本效率和有效性。

原文摘要 · Abstract (English)

Data valuation is a class of techniques for quantitatively assessing the value of data for applications like pricing in data marketplaces. Existing data valuation methods define a value for a discrete dataset. However, in many use cases, users are interested in not only the value of the dataset, but that of the distribution from which the dataset was sampled. For example, consider a buyer trying to evaluate whether to purchase data from different vendors. The buyer may observe (and compare) only a small preview sample from each vendor, to decide which vendor's data distribution is most useful to the buyer and purchase. The core question is how should we compare the values of data distributions from their samples? Under a Huber characterization of the data heterogeneity across vendors, we propose a maximum mean discrepancy (MMD)-based valuation method which enables theoretically principled and actionable policies for comparing data distributions from samples. We empirically demonstrate that our method is sample-efficient and effective in identifying valuable data distributions against several existing baselines, on multiple real-world datasets (e.g., network intrusion detection, credit card fraud detection) and downstream applications (classification, regression).

数据估值分布比较MMD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。