针对自诱导分布下的高斯过程主动学习,提出高效近似方法。
Active Learning for Gaussian Process Regression Under Self-Induced Boltzmann Weights

- 基于高斯过程设计新采集函数,闭式逼近未知分布
- 理论保证预测误差随样本量趋于零,且平均表现更优
- 适用于化学势能面与药物发现等实际场景
我们研究主动学习问题,目标是在函数自身诱导的未知玻尔兹曼分布下,以低预测误差学习未知函数。该自诱导加权在计算化学中的势能面建模中自然出现,但因目标分布未知且配分函数不可计算而带来独特挑战。本文提出 exttt{AB-SID-iVAR},一种基于高斯过程的采集函数,可在闭式中近似不可计算的贝叶斯目标分布,无需估计配分函数,适用于离散与连续输入域。同时分析了泰普森采样变体 exttt{TS-SID-iVAR},作为高方差蒙特卡洛版本。在温和条件下,证明终端预测误差以高概率趋于零,并给出更紧的平均情况保证。在合成基准及真实世界的势能面建模与药物发现任务中,均展现出对现有方法的一致改进。
原文摘要 · Abstract (English)
We consider the active learning problem where the goal is to learn an unknown function with low prediction error under an unknown Boltzmann distribution induced by the function itself. This self-induced weighting arises naturally in problems such as potential energy surface (PES) modeling in computational chemistry, yet poses unique challenges as the target distribution is unknown and its partition function is intractable. We propose \texttt{AB-SID-iVAR}, a Gaussian Process-based acquisition function that approximates the intractable Bayesian target distribution in closed form while avoiding partition function estimation, and is applicable to both discrete and continuous input domains. We also analyze a Thompson sampling alternative (\texttt{TS-SID-iVAR}) as a higher variance Monte Carlo variant. Despite the unknown target, under mild conditions, we establish that the terminal prediction error vanishes with high probability, and provide a tighter average-case guarantee. We demonstrate consistent improvements over existing approaches in this setting on synthetic benchmarks and real-world PES modeling and drug discovery tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。