arXiv:2606.21260stat.MLcs.LG2026-06

提出一种在核空间中优化采样的方法,降低计算成本同时保持精度。

Subsampling for supervised learning in reproducing kernel Hilbert spaces

论文配图:Subsampling for supervised learning in reproducing kernel Hilbert spaces
图 1 · 摘自论文原文
  • 基于霍维茨-汤普森加权重估损失,在再生核希尔伯特空间中设计采样策略。
  • 理论证明最优采样方案可使协方差算子迹最小,提升估计效率。
  • 实验证明该方法在合成与真实数据上均有效,适合大数据下高效学习场景。

在大数据时代,采样已成为统计学习中的常见实践。通过选择部分样本训练模型,采样旨在降低估计步骤的计算成本与时间,理想情况下还能减少能耗和碳足迹。本文研究非参数设定,假设假设空间位于再生核希尔伯特空间(reproducing kernel Hilbert space),且估计器为经霍维茨-汤普森(Horvitz-Thompson)重加权的经验风险最小化器。通过分析该估计器的渐近性质,揭示了以协方差算子迹最小化为目标的最优采样方案,并证明可通过插值法实现。在合成数据与真实数据集上的数值实验表明,该方法具有可行性与显著优势。

原文摘要 · Abstract (English)

In the era of big data, subsampling became a common practice in statistical learning. By selecting a subgroup of individuals based on which the learner is trained, subsampling aims at reducing the computational cost and time of the estimation step, and ideally leads to a decrease of its energy consumption and carbon footprint. This work focuses on a nonparametric setting, in which the hypotheses set lies in a reproducing kernel Hilbert space, and the estimator is a minimizer of an empirical risk reweighted à la Horvitz-Thompson. By studying the asymptotic properties of this estimator, we reveal an optimal subsampling scheme (regarding the trace of the covariance operator) and show that it can be used via plug-in. A numerical study on synthetic and real-world datasets shows the practicability and the benefit of the proposed approach.

采样优化核方法高效学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。