arXiv:2410.09940cs.LGcs.AI2024-10被引 5

将数据点分组进行影响度分析,速度提升10到50倍。

Generalized Group Data Attribution

  • 按数据组而非单个样本计算影响度,降低计算开销。
  • 在保持效果的前提下,速度比传统方法快10至50倍。
  • 适合需要大规模数据清洗与模型可解释性的研究者使用。

数据影响度(Data Attribution, DA)方法用于量化训练数据对模型输出的影响,广泛应用于可解释性、数据筛选和噪声标签识别。然而,现有方法通常计算成本高,难以应用于大规模机器学习模型。为此,我们提出广义分组数据影响度(GGDA)框架,通过将影响度归因于训练数据的分组而非单个数据点,显著简化计算过程。GGDA是一个通用框架,涵盖已有方法并支持未来新方法的集成,允许用户根据需求在效率与精度之间灵活权衡。实验表明,将GGDA应用于影响力函数、TracIn和TRAK等主流方法,相比标准方法实现最高达10至50倍的速度提升,同时维持良好的归因质量。在数据集剪枝和噪声标签识别等下游任务中,GGDA大幅提高计算效率并保持有效性,使此前难以实现的大规模应用场景变为可能。

原文摘要 · Abstract (English)

Data Attribution (DA) methods quantify the influence of individual training data points on model outputs and have broad applications such as explainability, data selection, and noisy label identification. However, existing DA methods are often computationally intensive, limiting their applicability to large-scale machine learning models. To address this challenge, we introduce the Generalized Group Data Attribution (GGDA) framework, which computationally simplifies DA by attributing to groups of training points instead of individual ones. GGDA is a general framework that subsumes existing attribution methods and can be applied to new DA techniques as they emerge. It allows users to optimize the trade-off between efficiency and fidelity based on their needs. Our empirical results demonstrate that GGDA applied to popular DA methods such as Influence Functions, TracIn, and TRAK results in upto 10x-50x speedups over standard DA methods while gracefully trading off attribution fidelity. For downstream applications such as dataset pruning and noisy label identification, we demonstrate that GGDA significantly improves computational efficiency and maintains effectiveness, enabling practical applications in large-scale machine learning scenarios that were previously infeasible.

数据影响度高效计算数据清洗可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。