条件核岭回归可提升预测性能,尤其当主效应明显时。
Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding

- 将线性回归与核方法结合,先拟合显著特征,再对残差做核回归。
- 理论证明其误差比标准核岭回归多一个1/√N量级的额外项。
- 适合处理主效应突出、可解释性强的回归任务。
条件正定核(CPD)相对于函数类 𝒫 定义,其对应的原生空间可导出一种学习方法——条件核岭回归(conditional KRR),该方法通过原生空间范数平方对回归函数进行惩罚。此方法可视为经典线性回归(以 𝒫 为特征)后接标准核岭回归(KRR)对残差部分建模。本文通过将条件KRR的性能等价为另一个固定核上的标准KRR,分析其统计性质。主要理论结果表明,这种等价成立,代价是期望测试风险增加一项,其上界为 𝒪(1/√N),其中 N 为样本量,常数依赖于 𝒫 和输入分布。我们研究了两种情形:一是 𝒫 由核 K 的前 k 个梅尔瑟特征函数构成;二是 𝒫 由随机特征表示中的随机选取的 k 个特征组成。二者关系密切。理论与实验均表明,当 𝒫 成分在回归函数中占主导时,条件KRR优于标准KRR。
原文摘要 · Abstract (English)
Conditionally positive definite (CPD) kernels are defined with respect to a function class $\mathcal{F}$. It is well known that such a kernel $K$ is associated with its native space (defined analogously to an RKHS), which in turn gives rise to a learning method -- called conditional kernel ridge regression (conditional KRR) due to its analogy with KRR -- where the estimated regression function is penalized by the square of its native space norm. This method is of interest because it can be viewed as classical linear regression, with features specified by $\mathcal{F}$, followed by the application of standard KRR to the residual (unexplained) component of the target variable. Methods of this type have recently attracted increasing attention. We study the statistical properties of this method by reducing its behavior to that of KRR with another fixed kernel, called the residual kernel. Our main theoretical result shows that such a reduction is indeed possible, at the cost of an additional term in the expected test risk, bounded by $\mathcal{O}(1/\sqrt{N})$, where $N$ is the sample size and the hidden constant depends on the class $\mathcal{F}$ and the input distribution. This reduction enables us to analyze conditional KRR in the case where $K$ is positive definite and $\mathcal{F}$ is given by the first $k$ principal eigenfunctions in the Mercer decomposition of $K$. We also consider the setting where $\mathcal{F}$ consists of $k$ random features from a random feature representation of $K$. It turns out that these two settings are closely related. Both our theoretical analysis and experiments confirm that conditional KRR outperforms standard KRR in these cases whenever the $\mathcal{F}$-component of the regression function is more pronounced than the residual part.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。