提出新型随机特征方法,高效逼近非平稳核函数。
Bernstein-Schur Kernels: Random Features by Sketched Modulation and Radial Randomization
- 通过调制与径向随机化联合构造随机特征,避免传统方法局限。
- 理论保证特征维度仅需O((1+‖P‖_op/λ)log(d_eff/δ)),显著降低计算开销。
- 适用于高维非平稳核,尤其适合机器学习中复杂相似度建模场景。
Bernstein-Schur 核是有限特征核与完全单调平移不变核的乘积,属于介于平移不变与点积模板之间的非平稳核,无法直接使用 Bochner 采样或多项式压缩。本文提出一种统一的随机特征构造方法:对有限调制部分进行随机化,同时在高斯随机傅里叶特征前对径向因子的 Bernstein-Widder 尺度进行一维采样,实现特征维度 $Dm$,且无需精确调制特征的 $O(d^2)$ 大小。当调制保持精确($m\to\infty$ 极限)时,证明了无偏性、精确方差及受顶层核与调制特征值控制的矩阵 Bernstein 模范界,其内在维度而非粗略的 $N\max_{ij}$ 路径。在岭回归下白化处理后,有效维度 $d_{\mathrm{eff}}(\lambda)$ 成为矩阵方差的精确内在维度,仅需 $O((1+\|P\|_{\mathrm{op}}/\lambda)\log(d_{\mathrm{eff}}/\delta))$ 次径向采样即可保留核岭解;通过闭式白化杠杆倾斜采样,进一步降至 $O((1+d_{\mathrm{eff}})\log(d_{\mathrm{eff}}/\delta))$。条件于采样结果的所有保证均能传递至双重随机估计器,仅增加一个额外采样项,且对整个类成立,只需用调制 Gram 矩阵替代多项式版本。旗舰实例为有偏 $yat$-核 $k_{yat,b}(w,x)=(w^\top x+b)^2/(\|w-x\|^2+\varepsilon)$,其家族包含由 $b$ 的有限差分生成的逆多二次核。
原文摘要 · Abstract (English)
Bernstein--Schur kernels are products of a finite-feature kernel and a completely monotone shift-invariant kernel: nonstationary kernels falling between the shift-invariant and dot-product templates random features exploit, so neither Bochner sampling nor polynomial sketching applies to the full kernel directly. We give one random-feature construction for the whole class that randomizes both factors: it sketches the finite modulation and samples the radial factor's one-dimensional Bernstein--Widder scale before applying Gaussian random Fourier features, giving feature dimension $Dm$, free of the $O(d^2)$ size of the exact modulation feature. With the modulation kept exact (the $m\to\infty$ limit), we prove unbiasedness, an exact variance, and a matrix-Bernstein operator-norm bound controlled by the top kernel and modulation eigenvalues and an intrinsic dimension rather than the crude $N\max_{ij}$ route. Whitening this argument at the ridge makes the effective dimension $d_{\mathrm{eff}}(λ)$ the \emph{exact} intrinsic dimension of the matrix variance, so $O((1+\|P\|_{\mathrm{op}}/λ)\log(d_{\mathrm{eff}}/δ))$ radial draws preserve the kernel-ridge solution; tilting the draw by a closed-form whitened leverage improves this to the effective-dimension count $O((1+d_{\mathrm{eff}})\log(d_{\mathrm{eff}}/δ))$. Conditioning on the sketch carries every guarantee to the deployed doubly-randomized estimator up to one additive sketch term, and all hold for the whole class with the modulation Gram in place of the polynomial one. The flagship instance is the biased $yat$-kernel $k_{yat,b}(w,x)=(w^\top x+b)^2/(\|w-x\|^2+\varepsilon)$, whose family span contains the inverse-multiquadric kernel by finite differences in $b$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。