从数据角度优化学习性能,用少量精选数据达到全量训练效果
Data Selection for ERMs
- 固定学习算法,通过筛选数据提升模型表现
- 在均值估计、线性分类与回归中实现最优数据选择边界
- 适用于关注数据效率的研究者和实际部署场景
学习理论传统上以模型为中心,专注于为固定自然学习任务(如线性分类或回归)设计最优算法。本文采用互补的数据中心视角,固定自然学习规则,聚焦于优化训练数据。具体而言,我们研究如下问题:给定学习规则 $\/mathcal{A}$ 与数据选择预算 $n$,当从 $N$ 个点的总体中最多选取 $n$ 个点时,$\/mathcal{A}$ 能达到何种性能?我们探讨在 $n \ll N$ 的情况下,能否实现与全量数据训练相当的性能。研究覆盖多种经验风险最小化器,得到均值估计、线性分类与线性回归的最优数据选择界限。此外,我们建立了二分类与随机凸优化中的误差率分类体系。最后提出若干开放问题与未来方向。
原文摘要 · Abstract (English)
Learning theory has traditionally followed a model-centric approach, focusing on designing optimal algorithms for a fixed natural learning task (e.g., linear classification or regression). In this paper, we adopt a complementary data-centric perspective, whereby we fix a natural learning rule and focus on optimizing the training data. Specifically, we study the following question: given a learning rule $\mathcal{A}$ and a data selection budget $n$, how well can $\mathcal{A}$ perform when trained on at most $n$ data points selected from a population of $N$ points? We investigate when it is possible to select $n \ll N$ points and achieve performance comparable to training on the entire population. We address this question across a variety of empirical risk minimizers. Our results include optimal data-selection bounds for mean estimation, linear classification, and linear regression. Additionally, we establish two general results: a taxonomy of error rates in binary classification and in stochastic convex optimization. Finally, we propose several open questions and directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。