提出鲁棒主动学习方法,让高斯过程回归在不确定分布下也能准确预测。
Distributionally Robust Active Learning for Gaussian Process Regression
- 基于最坏情况期望误差设计主动学习策略
- 理论上证明有限数据下误差可趋近于零
- 适合分布不确定场景下的非线性预测任务
高斯过程回归(GPR)或核岭回归是广泛使用的非线性预测工具。因此,针对GPR的主动学习(AL)——通过主动收集标签以较少样本实现高精度预测——是一个重要问题。然而,现有AL方法无法在目标分布上理论保证预测精度。此外,如分布鲁棒学习文献所指出,指定目标分布往往困难。本文提出两种AL方法,有效降低GPR的最坏情况期望误差,即在目标分布候选集中的最大期望误差。我们推导了最坏情况期望平方误差的上界,表明在温和条件下,通过有限数量的数据标签,误差可任意小。最后,我们在合成与真实世界数据集上验证了所提方法的有效性。
原文摘要 · Abstract (English)
Gaussian process regression (GPR) or kernel ridge regression is a widely used and powerful tool for nonlinear prediction. Therefore, active learning (AL) for GPR, which actively collects data labels to achieve an accurate prediction with fewer data labels, is an important problem. However, existing AL methods do not theoretically guarantee prediction accuracy for target distribution. Furthermore, as discussed in the distributionally robust learning literature, specifying the target distribution is often difficult. Thus, this paper proposes two AL methods that effectively reduce the worst-case expected error for GPR, which is the worst-case expectation in target distribution candidates. We show an upper bound of the worst-case expected squared error, which suggests that the error will be arbitrarily small by a finite number of data labels under mild conditions. Finally, we demonstrate the effectiveness of the proposed methods through synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。