用分布编码提升类别输入的高斯过程回归性能
Distributional encoding for Gaussian process regression with qualitative inputs
- 基于类别样本分布构建编码,而非仅用均值
- 在多个真实与合成数据集上达到顶尖预测效果
- 适合需要高效采样和离散变量优化的研究者
高斯过程回归在工程应用中广受欢迎,尤其适用于观测代价高昂的场景,也是贝叶斯优化的核心方法。然而,当输入变量包含或全为类别型时,构建高效且准确的高斯过程仍具挑战。本文从朴素目标编码出发,提出分布编码(DE),利用每个类别下所有目标值的完整分布信息,而非仅使用均值。为在高斯过程中处理此类编码,我们引入基于最大均值差异和沃尔什距离的特征核理论。此外,还讨论了分类、多任务学习及辅助信息融合等扩展。实验验证表明,该方法在多种合成与真实数据集上表现卓越,达到当前最优水平。该方法天然契合离散与混合空间贝叶斯优化的最新进展。
原文摘要 · Abstract (English)
Gaussian Process (GP) regression is a popular and sample-efficient approach for many engineering applications, where observations are expensive to acquire, and is also a central ingredient of Bayesian optimization (BO), a highly prevailing method for the optimization of black-box functions. However, when all or some input variables are categorical, building a predictive and computationally efficient GP remains challenging. Starting from the naive target encoding idea, where the original categorical values are replaced with the mean of the target variable for that category, we propose a generalization based on distributional encoding (DE) which makes use of all samples of the target variable for a category. To handle this type of encoding inside the GP, we build upon recent results on characteristic kernels for probability distributions, based on the maximum mean discrepancy and the Wasserstein distance. We also discuss several extensions for classification, multi-task learning and incorporation or auxiliary information. Our approach is validated empirically, and we demonstrate state-of-the-art predictive performance on a variety of synthetic and real-world datasets. DE is naturally complementary to recent advances in BO over discrete and mixed-spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。