提出高效子数据选择方法,显著提升小样本下模型估计精度。
Nearly Optimal Subdata Selection
- 基于最优近似设计理论,构建通用子数据选择算法。
- 所选子数据效率逼近理论最优,超越现有所有方法。
- 可评估任意方法的子数据效率,适用于大规模或高成本场景。
当数据集规模N超过计算资源限制或标注成本过高时,仅选取部分数据点(子数据)进行分析成为有效方案。核心问题是:从N个数据点中选出n个最优子数据。针对参数化模型与参数估计目标,理想策略是保留最大信息量的子数据。由于该问题具有离散性,属于经典NP难问题。本文基于最优近似设计理论,提出一种新的信息导向子数据选择方法,所选子数据可逼近理论最优解。为此,我们开发了一种新型算法,适用于一般模型、任意的N和n值,并支持多种最优性准则,且证明了其收敛性。此外,该方法可为任意子数据选择方法提供效率的紧致上下界评估。实验表明,新方法所得子数据效率极高,全面优于现有方法。
原文摘要 · Abstract (English)
When, in terms of the number of data points, the size of a dataset exceeds available computing resources, or when labeling is expensive, an attractive solution consists of selecting only some of the data points (subdata) for further consideration. A central question for selecting subdata of size $n$ from $N$ available data points is which $n$ points to select. While an answer to this question depends on the objective, one approach for a parametric model and a focus on parameter estimation is to select subdata that retains maximal information. Identifying such subdata is a classical NP-hard problem due to its inherent discreteness. Based on optimal approximate design theory, we develop a new methodology for information-based subdata selection, resulting in subdata that approaches the optimal solution. To achieve this, we develop a novel algorithm that applies to a general model, accommodates arbitrary choices of $N$ and $n$, and supports multiple optimality criteria, and we prove its convergence. Moreover, the new methodology facilitates an assessment of the efficiency of subdata selected by any method by obtaining tight lower and upper bounds for the efficiency. We show that the subdata obtained through the new methodology is highly efficient and outperforms all existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。