arXiv:2410.02680stat.MLcs.LG2024-10被引 3

提出一种自适应岭回归,对表格数据建模效果更优。

Highly Adaptive Ridge

  • 基于自适应核的核岭回归,使用饱和零阶张量样条展开。
  • 在右连续函数类中实现 $n^{-1/3}$ 的无维度误差收敛率。
  • 小数据集上表现超越现有算法,适合表格数据建模。

本文提出高度自适应岭(HAR):一种回归方法,在右连续且具有平方可积截面导数的函数类中实现了 $n^{-1/3}$ 的无维度 L2 收敛速率。该非参数函数类特别适用于表格数据。HAR 是一种特定数据自适应核的核岭回归,其核基于饱和零阶张量积样条基展开。通过模拟和真实数据验证了理论结果。实证显示,在小数据集上性能优于当前最优算法。

原文摘要 · Abstract (English)

In this paper we propose the Highly Adaptive Ridge (HAR): a regression method that achieves a $n^{-1/3}$ dimension-free L2 convergence rate in the class of right-continuous functions with square-integrable sectional derivatives. This is a large nonparametric function class that is particularly appropriate for tabular data. HAR is exactly kernel ridge regression with a specific data-adaptive kernel based on a saturated zero-order tensor-product spline basis expansion. We use simulation and real data to confirm our theory. We demonstrate empirical performance better than state-of-the-art algorithms for small datasets in particular.

回归自适应表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。