提出一种确定性核密度压缩方法,可精准逼近机器学习优化目标。
Deterministic Coreset Construction via Adaptive Sensitivity Trimming
- 通过自适应剪枝低敏感度数据点并均匀加权剩余点构造核心集。
- 理论保证对所有假设空间实现(1±ε)相对误差的统一近似。
- 适用于核岭回归等模型,支持复现且可扩展至流式与公平约束场景。
我们构建了一个严格的确定性核心集构造框架,用于经验风险最小化(ERM)。核心贡献是自适应确定性均匀权重剪枝(ADUWT)算法:通过剔除敏感度下界最低的数据点,并对剩余数据赋予依赖数据的均匀权重,构造出核心集。该方法在全假设空间上实现了对ERM目标的(1±ε)相对误差统一近似。我们提供了完整分析,包括:(i) 最小最大表征证明自适应权重的最优性;(ii) 基于敏感度异质指数的实例相关大小分析;(iii) 核岭回归、正则化逻辑回归和线性SVM的可计算敏感度预言机。算法伪代码、敏感度预言机及评估流程均明确给出,支持复现。实验结果与理论一致。最后,我们提出了关于实例最优预言机、确定性流式处理及公平约束ERM的开放问题。
原文摘要 · Abstract (English)
We develop a rigorous framework for deterministic coreset construction in empirical risk minimization (ERM). Our central contribution is the Adaptive Deterministic Uniform-Weight Trimming (ADUWT) algorithm, which constructs a coreset by excising points with the lowest sensitivity bounds and applying a data-dependent uniform weight to the remainder. The method yields a uniform $(1\pm\varepsilon)$ relative-error approximation for the ERM objective over the entire hypothesis space. We provide complete analysis, including (i) a minimax characterization proving the optimality of the adaptive weight, (ii) an instance-dependent size analysis in terms of a \emph{Sensitivity Heterogeneity Index}, and (iii) tractable sensitivity oracles for kernel ridge regression, regularized logistic regression, and linear SVM. Reproducibility is supported by precise pseudocode for the algorithm, sensitivity oracles, and evaluation pipeline. Empirical results align with the theory. We conclude with open problems on instance-optimal oracles, deterministic streaming, and fairness-constrained ERM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。