arXiv:2505.09134cs.LGstat.ML2025-05被引 1

提出可扩展的高斯过程方法,支持完整导数观测的高效建模。

Scaling Gaussian Process Regression with Full Derivative Observations

  • 用局部温度向量替代全局温度,提升导数敏感度建模能力。
  • 在100-1000维分子力场预测中表现准确,支持大规模数据。
  • 无需显式核函数导数,适合深度核学习等扩展应用。

我们提出一种可扩展的高斯过程方法 DSoftKI,能够拟合并预测完整的导数观测数据。该方法将 SoftKI(通过 softmax 插值近似核函数)扩展至导数场景,通过将全局温度向量替换为与每个插值点关联的局部温度向量,增强模型对局部方向敏感性的编码能力。这一改进使我们能够通过插值构建包含一阶和二阶导数的可扩展近似核。此外,该插值方案避免了对核函数导数的显式计算,从而便于扩展至深度核学习(DKL)。我们在合成基准、简化的多体物理模拟、含合成梯度的标准回归数据集以及高维分子力场预测(100-1000维)上评估了 DSoftKI。结果表明,DSoftKI 在精度与规模上均优于以往方法,首次实现了在完整导数观测下对大规模数据的有效建模。

原文摘要 · Abstract (English)

We present a scalable Gaussian Process (GP) method called DSoftKI that can fit and predict full derivative observations. It extends SoftKI, a method that approximates a kernel via softmax interpolation, to the setting with derivatives. DSoftKI enhances SoftKI's interpolation scheme by replacing its global temperature vector with local temperature vectors associated with each interpolation point. This modification allows the model to encode local directional sensitivity, enabling the construction of a scalable approximate kernel, including its first and second-order derivatives, through interpolation. Moreover, the interpolation scheme eliminates the need for kernel derivatives, facilitating extensions such as Deep Kernel Learning (DKL). We evaluate DSoftKI on synthetic benchmarks, a toy n-body physics simulation, standard regression datasets with synthetic gradients, and high-dimensional molecular force field prediction (100-1000 dimensions). Our results demonstrate that DSoftKI is accurate and scales to larger datasets with full derivative observations than previously possible.

高斯过程导数观测可扩展性核方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。