arXiv:2608.26974stat.MLcs.LG2026-08

高斯核导致预测过度自信,应避免作为默认选择

Why not to use the Gaussian kernel

  • 高斯核导致条件方差过小,引发预测不确定性被严重低估
  • 其数值不稳定性需引入人为修正项,改变原始模型设定
  • 所有解析核都具类似问题,应谨慎选用此类光滑核

核函数广泛用于回归与分类任务中衡量相似性。高斯核(又称平方指数核、径向基函数核)在高斯过程回归中极为流行。本文主张应避免使用高斯核,且不应将其作为默认选择。论据基于两项发现:其一,高斯核导致的条件方差异常偏小,若以此量化预测不确定性,将几乎必然引发灾难性的过度自信;其二,极小方差伴随严重的数值病态,实际应用中必须引入像噪声项(nugget term)之类的技巧,实质上修改了原始回归或分类模型。这些问题源于高斯核的非自然平滑性——并非其高斯形式本身的问题,而是其解析性所致。更广泛的结论是:解析核应尽量避免。对于平稳核而言,解析性等价于谱密度呈指数衰减。

原文摘要 · Abstract (English)

Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. We argue that the Gaussian kernel is best avoided and should never be used as a default. The argument rests on two results demonstrating that the Gaussian kernel is extremely brittle. First, the Gaussian kernel gives rise to a conditional variance that is unrealistically small. If the variance is used to quantify predictive uncertainty, catastrophic overconfidence is almost inevitable. Second, a small variance goes hand in hand with numerical ill-conditioning, so that to use the Gaussian kernel in practice requires tricks such as nugget terms that effectively modify the underlying regression or classification model. These problems are caused by the unnatural smoothness of the Gaussian kernel, a fact we are far from the first to take notice of. The problem is not the Gaussian form itself but the analyticity of the kernel: Our argument is more broadly that analytic kernels are best avoided. For stationary kernels analyticity is essentially equivalent to an exponential decay of the spectral density.

高斯过程核方法预测不确定数值稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。