arXiv:2409.18153cs.LGstat.ML2024-09NeurIPS被引 37

找出对模型影响最大的训练数据子集,揭示集体效应的复杂性。

Most Influential Subset Selection: Challenges, Promises, and Beyond

  • 用自适应贪心算法捕捉样本间的集体影响
  • 理论证明传统方法在线性回归中可能失效
  • 适用于分类与非线性神经网络,兼顾效率与效果

如何将机器学习模型的行为归因于其训练数据?经典影响函数虽能分析单个样本的影响,却难以捕捉样本集合的复杂集体效应。为此,本文研究最具有影响力的子集选择(MISS)问题,旨在识别具有最大集体影响的训练样本子集。我们全面分析了现有方法的优劣,发现基于影响函数的贪心启发式算法在某些情况下会明显失效,根源在于影响函数误差及集体影响的非加性结构。相反,迭代应用此类启发式的自适应版本能有效建模样本间交互,部分缓解上述问题。真实数据集上的实验验证了理论结论,并表明自适应策略在分类任务和非线性神经网络中仍具优势。最后,我们强调性能与计算效率间的权衡,质疑加性度量如线性数据建模得分的适用性,并展开多方面讨论。

原文摘要 · Abstract (English)

How can we attribute the behaviors of machine learning models to their training data? While the classic influence function sheds light on the impact of individual samples, it often fails to capture the more complex and pronounced collective influence of a set of samples. To tackle this challenge, we study the Most Influential Subset Selection (MISS) problem, which aims to identify a subset of training samples with the greatest collective influence. We conduct a comprehensive analysis of the prevailing approaches in MISS, elucidating their strengths and weaknesses. Our findings reveal that influence-based greedy heuristics, a dominant class of algorithms in MISS, can provably fail even in linear regression. We delineate the failure modes, including the errors of influence function and the non-additive structure of the collective influence. Conversely, we demonstrate that an adaptive version of these heuristics which applies them iteratively, can effectively capture the interactions among samples and thus partially address the issues. Experiments on real-world datasets corroborate these theoretical findings and further demonstrate that the merit of adaptivity can extend to more complex scenarios such as classification tasks and non-linear neural networks. We conclude our analysis by emphasizing the inherent trade-off between performance and computational efficiency, questioning the use of additive metrics such as the Linear Datamodeling Score, and offering a range of discussions.

数据归因影响函数集合选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。