arXiv:2511.02100cs.LGcs.AI2025-11

用几何杠杆得分快速评估数据重要性,媲美经典方法且更高效

Geometric Data Valuation via Leverage Scores

  • 基于数据在表示空间的结构影响,用杠杆得分替代复杂组合计算
  • 采样杠杆得分数据可使模型性能接近全量数据最优,误差在O(ε)内
  • 无需梯度即可实现优于基线的主动学习,适合大规模数据筛选

Shapley数据估值提供了一种严谨的、公理化的数据点重要性分配框架,在数据集整理、剪枝和定价中日益受到关注。然而,其组合性质需对所有数据子集计算边际贡献,难以规模化。本文提出一种基于统计杠杆得分的几何替代方案,通过衡量每个数据点在表示空间中扩展数据集跨度及贡献有效维度的能力,量化其结构性影响。我们证明该得分满足虚无、效率和对称性公理;进一步引入岭杠杆得分,可获得严格正的边际收益,并自然关联经典A-和D-最优设计准则。我们还证明,基于杠杆得分采样的子集训练出的模型,其参数与预测风险均在O(ε)范围内逼近全数据最优解,建立了数据估值与下游决策质量的严格联系。最后,在主动学习实验中,我们实证表明岭杠杆采样优于标准基线,且无需梯度或反向传播。

原文摘要 · Abstract (English)

Shapley data valuation provides a principled, axiomatic framework for assigning importance to individual datapoints, and has gained traction in dataset curation, pruning, and pricing. However, it is a combinatorial measure that requires evaluating marginal utility across all subsets of the data, making it computationally infeasible at scale. We propose a geometric alternative based on statistical leverage scores, which quantify each datapoint's structural influence in the representation space by measuring how much it extends the span of the dataset and contributes to the effective dimensionality of the training problem. We show that our scores satisfy the dummy, efficiency, and symmetry axioms of Shapley valuation and that extending them to \emph{ridge leverage scores} yields strictly positive marginal gains that connect naturally to classical A- and D-optimal design criteria. We further show that training on a leverage-sampled subset produces a model whose parameters and predictive risk are within $O(\varepsilon)$ of the full-data optimum, thereby providing a rigorous link between data valuation and downstream decision quality. Finally, we conduct an active learning experiment in which we empirically demonstrate that ridge-leverage sampling outperforms standard baselines without requiring access gradients or backward passes.

数据估值杠杆得分主动学习高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。