简单线性模型在高维贝叶斯优化中表现超越复杂方法。
We Still Don't Understand High-Dimensional Bayesian Optimization
- 用线性核高斯过程替代复杂模型,配合几何变换避免边界陷阱。
- 在60至6000维空间中达到顶尖性能,分子优化任务超2万样本仍高效。
- 计算线性增长,适合大规模数据,挑战高维优化传统认知。
现有高维贝叶斯优化方法通过引入局部性、稀疏性或光滑性等结构假设来缓解维度灾难。然而,我们发现这些方法反而被最简单的贝叶斯线性回归所超越。经过几何变换以避免边界偏好后,采用线性核的高斯过程在60至6000维搜索空间任务中达到当前最优表现。线性模型相比非参数方法具有显著优势:支持闭式采样,计算复杂度随数据线性增长,这一特性使其在包含超过20,000个观测值的分子优化任务中依然高效。结合实证分析,结果表明需重新审视高维贝叶斯优化的传统范式。
原文摘要 · Abstract (English)
Existing high-dimensional Bayesian optimization (BO) methods aim to overcome the curse of dimensionality by carefully encoding structural assumptions, from locality to sparsity to smoothness, into the optimization procedure. Surprisingly, we demonstrate that these approaches are outperformed by arguably the simplest method imaginable: Bayesian linear regression. After applying a geometric transformation to avoid boundary-seeking behavior, Gaussian processes with linear kernels match state-of-the-art performance on tasks with 60- to 6,000-dimensional search spaces. Linear models offer numerous advantages over their non-parametric counterparts: they afford closed-form sampling and their computation scales linearly with data, a fact we exploit on molecular optimization tasks with >20,000 observations. Coupled with empirical analyses, our results suggest the need to depart from past intuitions about BO methods in high-dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。