arXiv:2512.12731cs.LGcs.NA2025-12

基于随机函数理论推导出最优回归方法,无需经验选择核函数。

Solving a Machine Learning Regression Problem Based on the Theory of Random Functions

  • 从对称性假设出发,数学推导出回归核函数与正则化形式。
  • 所得核函数为广义多调和样条,且在无先验信息下最优。
  • 适合追求理论严谨性的机器学习研究者参考。

本文将机器学习回归问题视为多元逼近问题,基于随机函数理论框架进行研究。提出一种从对称性假设(平移、旋转、缩放不变性及高斯性)出发的回归方法推导过程。若无穷维函数空间上的概率测度具备这些自然对称性,则整个求解方案——包括核函数形式、正则化类型和噪声参数化方式——可由这些假设严格解析得出。所得到的核函数与广义多调和样条一致,但并非经验选定,而是源于‘无关性原理’的自然结果。该结论为一大类光滑化与插值方法提供了理论基础,证明了其在缺乏先验信息时的最优性。

原文摘要 · Abstract (English)

This paper studies a machine learning regression problem as a multivariate approximation problem using the framework of the theory of random functions. An ab initio derivation of a regression method is proposed, starting from postulates of indifference. It is shown that if a probability measure on an infinite-dimensional function space possesses natural symmetries (invariance under translation, rotation, scaling, and Gaussianity), then the entire solution scheme, including the kernel form, the type of regularization, and the noise parameterization, follows analytically from these postulates. The resulting kernel coincides with a generalized polyharmonic spline; however, unlike existing approaches, it is not chosen empirically but arises as a consequence of the indifference principle. This result provides a theoretical foundation for a broad class of smoothing and interpolation methods, demonstrating their optimality in the absence of a priori information.

回归分析理论推导核方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。