arXiv:2605.10684cs.LGcs.AI2026-05中稿 · ICML被引 1

提出NASH框架,让数据选择更有效且高效。

Is Data Shapley Not Better than Random in Data Selection? Ask NASH

论文配图:Is Data Shapley Not Better than Random in Data Selection? Ask NASH
图 1 · 摘自论文原文
  • 将目标函数分解为可解释的组件,利用非线性聚合提升选择效果
  • 在多个数据集上显著优于传统方法,且运行时间几乎不变
  • 适合需要高效高质量数据筛选的研究者和工程应用

数据选择旨在识别训练数据中高质量的子集。尽管部分研究采用高阶数据Shapley值或其它半值来考虑数据间的交互关系,但也有研究指出其在实际中表现不佳,选出的数据子集可能与随机选择无异。这引发两个问题:(I) 是否存在能让数据Shapley始终有效的‘有信息量’设置?(II) 能否利用这些设置实现稳定高效的高质量子集选择?本文提出一种新框架NASH(非线性聚合Shapley信息成分),(I) 将目标效用函数(如验证准确率)分解为更简单的Shapley信息组件,(II) 通过非线性聚合这些组件进行数据选择。实验表明,NASH在几乎不增加运行时间的前提下,显著提升了基于Shapley/半值的数据选择效果。

原文摘要 · Abstract (English)

Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top-$m$ Data Shapley or other semivalues as they account for the interaction among every subset of data, other works argue that Data Shapley can sometimes perform ineffectively in practice and select subsets that are no better than random. This raises the questions: (I) Are there certain "Shapley-informative" settings where Data Shapley consistently works well? (II) Can we strategically utilize these settings to select high-quality subsets consistently and efficiently? In this paper, we propose a novel data selection framework, NASH (Non-linear Aggregation of SHapley-informative components), which (I) decomposes the target utility function (e.g., validation accuracy) into simpler, Shapley-informative component functions, and selects data by optimizing an objective that (II) aggregates these components non-linearly. We demonstrate that NASH substantially boosts the effectiveness of Shapley/semivalue-based data selection with minimal additional runtime cost.

数据选择Shapley值高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。