系统梳理数据压缩三类方法,揭示其在模型训练中的新应用与挑战
A Coreset Selection of Coreset Selection Literature: Introduction and Recent Advances
- 按无训练、面向训练、无标签三类统一现有研究框架
- 提出子模优化与双层优化等被忽视的理论工具
- 适合关注高效训练与大模型压缩的研究者
核心集选择旨在从大规模数据集中筛选出能保留关键模式的小型代表性子集,以提升机器学习效率。尽管已有多个综述讨论过数据缩减策略,但多数仅聚焦于经典几何方法或主动学习技术。本文首次将训练无关、训练导向和无标签三类核心集研究统一为一个完整分类体系,涵盖子模优化、双层优化及伪标签在无标签数据上的最新进展。同时,分析剪枝策略对泛化能力与神经网络缩放定律的影响,填补了以往综述的空白。最后,在不同计算资源、鲁棒性与性能需求下对比各类方法,指出未来需解决的开放问题:鲁棒性、异常值过滤及适配基础模型的核心集选择。
原文摘要 · Abstract (English)
Coreset selection targets the challenge of finding a small, representative subset of a large dataset that preserves essential patterns for effective machine learning. Although several surveys have examined data reduction strategies before, most focus narrowly on either classical geometry-based methods or active learning techniques. In contrast, this survey presents a more comprehensive view by unifying three major lines of coreset research, namely, training-free, training-oriented, and label-free approaches, into a single taxonomy. We present subfields often overlooked by existing work, including submodular formulations, bilevel optimization, and recent progress in pseudo-labeling for unlabeled datasets. Additionally, we examine how pruning strategies influence generalization and neural scaling laws, offering new insights that are absent from prior reviews. Finally, we compare these methods under varying computational, robustness, and performance demands and highlight open challenges, such as robustness, outlier filtering, and adapting coreset selection to foundation models, for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。