arXiv:2506.04026cs.LG2025-06

用高斯过程快速评估数据对模型的影响,理论扎实且计算高效。

On the Usage of Gaussian Process for Efficient Data Valuation

  • 将数据估值分解为模型特性与信息聚合两部分,统一分析框架。
  • 利用高斯过程在子模型上快速估算数据价值,支持高效更新。
  • 适合关注数据质量评估、模型可解释性的研究人员使用。

在机器学习中,衡量单个样本对模型训练的影响是一项基础任务,称为数据估值。基于文献中的前期工作,我们设计了一种新的规范分解方法,使从业者能够将任何数据估值方法视为两个部分的组合:一个捕捉模型特征的效用函数,以及一个整合此类信息的聚合过程。我们提出使用高斯过程来便捷地获取在‘子模型’(即在训练集子集上训练的模型)上的效用函数。该方法的优势源于其在贝叶斯理论上的坚实基础,以及在实践中通过高效的更新公式实现快速估值的能力。

原文摘要 · Abstract (English)

In machine learning, knowing the impact of a given datum on model training is a fundamental task referred to as Data Valuation. Building on previous works from the literature, we have designed a novel canonical decomposition allowing practitioners to analyze any data valuation method as the combination of two parts: a utility function that captures characteristics from a given model and an aggregation procedure that merges such information. We also propose to use Gaussian Processes as a means to easily access the utility function on ``sub-models'', which are models trained on a subset of the training set. The strength of our approach stems from both its theoretical grounding in Bayesian theory, and its practical reach, by enabling fast estimation of valuations thanks to efficient update formulae.

数据估值高斯过程模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。