研究数据价值评估中效用变化对结果的影响,提出几何方法提升评估稳定性。
On the Impact of the Utility in Semivalue-based Data Valuation
- 将数据点嵌入低维空间,使效用变为线性函数,简化估值框架。
- 设计可量化的鲁棒性指标,明确评估结果受效用变动影响程度。
- 适用于多准则权衡场景,帮助选择更稳定的效用和半值方法。
基于半值的数据价值评估利用合作博弈论思想,为每个数据点分配反映其对下游任务贡献的数值。然而,这些价值依赖于实践者选择的效用函数,引发关键问题:当效用发生变化时,评估结果是否稳健?尤其在效用需权衡多个标准且存在多个等效选择时,该问题尤为突出。为此,本文引入数据集的“空间特征”概念:给定一个半值,将每个数据点嵌入低维空间,使得任意效用函数在此空间中均表现为线性泛函,从而让数据价值评估具备更简洁的几何解释。基于此,我们提出一种实用方法,核心是显式的鲁棒性度量,指导实践者判断其数据价值评估结果是否会随效用变化而显著偏移。我们在多种数据集和半值方法上验证了该方法,结果显示与秩相关分析高度一致,并揭示了选择特定半值如何放大或削弱评估的鲁棒性。
原文摘要 · Abstract (English)
Semivalue-based data valuation uses cooperative-game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the practitioner's choice of utility, raising the question: How robust is semivalue-based data valuation to changes in the utility? This issue is critical when the utility is set as a trade-off between several criteria and when practitioners must select among multiple equally valid utilities. We address this by introducing the notion of a dataset's spatial signature: given a semivalue, we embed each data point into a lower-dimensional space in which any utility becomes a linear functional, making the data valuation framework amenable to a simpler geometric picture. Building on this, we propose a practical methodology centered on an explicit robustness metric that informs practitioners whether and by how much their data valuation results will shift as the utility changes. We validate this approach across diverse datasets and semivalues, demonstrating strong agreement with rank-correlation analyses and offering analytical insight into how choosing a semivalue can amplify or diminish robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。