arXiv:2510.18332stat.MLcs.LG2025-10

提出训练数据异质性参数,揭示其对模型学习的关键影响。

Parametrising the Inhomogeneity Inducing Capacity of a Training Set, and its Impact on Supervised Learning

  • 定义训练数据异质性参数,衡量模型所需非平稳相关结构的程度。
  • 实证表明该参数影响真实多变量函数预测的精度与可靠性。
  • 适用于关注数据分布特性对模型性能影响的研究者。

我们引入一种对训练数据集特性的参数化方法,该特性决定了所学函数需具备非均匀相关结构以建模变量间关系。我们将这种数据集属性的参数化称为「异质性参数」。该参数易于在中小规模数据集上计算,并在多个公开数据集上完成演示;同时证明传统意义上的数据非平稳性并不等同于非零异质性参数。在基于高斯过程的概率学习框架下,我们证明:若训练集具有非零异质性参数,则建模目标函数的过程必须是非平稳的。在真实世界多变量函数的学习中,测试输入处的预测质量与可靠性均受训练数据异质性参数的影响。

原文摘要 · Abstract (English)

We introduce parametrisation of that property of the available training dataset, that necessitates an inhomogeneous correlation structure for the function that is learnt as a model of the relationship between the pair of variables, observations of which comprise the considered training data. We refer to a parametrisation of this property of a given training set, as its ``inhomogeneity parameter''. It is easy to compute this parameter for small-to-large datasets, and we demonstrate such computation on multiple publicly-available datasets, while also demonstrating that conventional ``non-stationarity'' of data does not imply a non-zero inhomogeneity parameter of the dataset. We prove that - within the probabilistic Gaussian Process-based learning approach - a training set with a non-zero inhomogeneity parameter renders it imperative, that the process that is invoked to model the sought function, be non-stationary. Following the learning of a real-world multivariate function with such a Process, quality and reliability of predictions at test inputs, are demonstrated to be affected by the inhomogeneity parameter of the training data.

高斯过程非平稳性数据特性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。