arXiv:2505.14083stat.MLcs.LG2025-05NeurIPS被引 4

用随机投影降低核方法计算开销,保持预测精度。

Computational Efficiency under Covariate Shift in Kernel Ridge Regression

  • 在核希尔伯特空间中使用随机子空间近似
  • 即使输入分布不同,仍可大幅减少内存和时间消耗
  • 适合大规模数据下需兼顾效率与准确的场景

本文研究再生核希尔伯特空间(RKHS)中非参数回归的协变量偏移问题。当训练与测试数据的输入分布不一致时,协变量偏移会带来额外学习挑战。尽管核方法具有最优统计性质,但其在时间和内存上的高开销限制了其在大数据集上的可扩展性。本文聚焦于在协变量偏移下计算效率与统计准确性之间的权衡,探索使用随机投影构建随机子空间作为假设空间的可行性。结果表明,即便存在协变量偏移,该方法仍能在不牺牲学习性能的前提下实现显著的计算节省。

原文摘要 · Abstract (English)

This paper addresses the covariate shift problem in the context of nonparametric regression within reproducing kernel Hilbert spaces (RKHSs). Covariate shift arises in supervised learning when the input distributions of the training and test data differ, presenting additional challenges for learning. Although kernel methods have optimal statistical properties, their high computational demands in terms of time and, particularly, memory, limit their scalability to large datasets. To address this limitation, the main focus of this paper is to explore the trade-off between computational efficiency and statistical accuracy under covariate shift. We investigate the use of random projections where the hypothesis space consists of a random subspace within a given RKHS. Our results show that, even in the presence of covariate shift, significant computational savings can be achieved without compromising learning performance.

核方法协变量偏移随机投影计算效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。