通过可学习变换矩阵提升核回归效率
A Variational Analysis of Kernel Learning with Learnable Linear Transformations
- 引入可学习变换矩阵优化核回归的尺度与特征
- 在多尺度和低维特征场景下显著提升性能
- 适用于需要自动提取数据内在结构的任务
经典核岭回归旨在用输入数据 $X\in \mathbb{R}^d$ 拟合输出 $Y$,依赖于固定再生核希尔伯特空间中的正则化项,如 Sobolev 空间。本文提出一种推广形式,引入额外矩阵参数 $U$,用于探测数据中的尺度参数和特征变量,从而提高核岭回归效率。这自然引出一个非线性变分问题,需优化 $U$ 的选择。我们研究了该变分问题的基础数学性质,包括欧拉-拉格朗日方程、连续性与一阶变分、退化或发散变换下的极限行为,以及局部极小值的结构。特别关注两类数据分布设置:多尺度模型和多指标模型,其中学习到的变换 $U$ 分别编码了内在尺度参数和关键低维特征变量。
原文摘要 · Abstract (English)
The classical kernel ridge regression problem aims to find the best fit for the output $Y$ as a function of the input data $X\in \mathbb{R}^d$, with a fixed choice of regularization term imposed by a given choice of a reproducing kernel Hilbert space, such as a Sobolev space. Here we consider a generalization of the kernel ridge regression problem, by introducing an extra matrix parameter $U$, which aims to detect the scale parameters and the feature variables in the data, and thereby improve the efficiency of kernel ridge regression. This naturally leads to a nonlinear variational problem to optimize the choice of $U$. We study various foundational mathematical aspects of this variational problem, including its Euler-Lagrange equation, continuity and first variation, limiting behavior under degenerate or diverging transformations, and the structure of its local minimizers. Particular attention is given to two data-distribution settings, namely multi-scale and multi-index models, where the learned transformation $U$ encodes intrinsic scale parameters and the essential low-dimensional feature variables, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。