针对数据稀缺场景,提出新方法优化数据删减效果
Constraint-Data-Value-Maximization: Utilizing Data Attribution for Effective Data Pruning in Low-Data Environments

- 将删减数据建模为带约束的优化问题,兼顾整体影响与单样本贡献
- 在仅保留少量数据时仍保持模型性能稳定,优于传统方法
- 适合数据量少但需高效筛选的机器学习应用
将模型行为归因于训练数据是当前研究热点。常用基准为数据移除:剔除价值低或高的数据实例后评估模型性能。现有研究多采用基于Shapley的数据值进行此任务。本文指出,在仅剩少量数据时,这些数据值并非最优选择。为此,我们提出约束-数据值最大化(CDVM)方法,有效利用数据归因实现低数据环境下的数据删减。通过将删减建模为同时最大化总影响力并惩罚单个测试样本过量贡献的约束优化问题,CDVM 在仅保留小部分数据时仍表现稳健。在OpenDataVal基准上,CDVM展现出优异性能和有竞争力的运行效率。
原文摘要 · Abstract (English)
Attributing model behavior to training data is an evolving research field. A common benchmark is data removal, which involves eliminating data instances with either low or high values, then assessing a model's performance trained on the modified dataset. Many existing studies leverage Shapley-based data values for this task. In this paper, we demonstrate that these data values are not optimally suited for pruning low-value data when only a limited amount of data remains. To address this limitation, we introduce the Constraint-Data-Value-Maximization (CDVM) approach, which effectively utilizes data attributions for pruning in low-data scenarios. By casting pruning as a constrained optimization that both maximizes total influence and penalizes excessive per-test contributions, CDVM delivers robust performance when only a small fraction of the data is retained. On the OpenDataVal benchmark, CDVM shows strong performance and competitive runtime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。