arXiv:2608.04113cs.LGcs.AI2026-08

利用历史高保真数据提升昂贵优化效率

Out-Of-The-Loop Multi-Fidelity Bayesian Optimization

论文配图:Out-Of-The-Loop Multi-Fidelity Bayesian Optimization
图 1 · 摘自论文原文
  • 引入历史高保真数据与任务描述符改进多保真贝叶斯优化
  • 在化学与超参优化中显著优于传统方法
  • 适合有历史实验数据的科研与工程场景

黑箱优化在科学与工程中普遍存在,常面临目标函数昂贵而低保真代理函数较便宜的问题。多保真贝叶斯优化(MF-BO)通过不同保真度间的相关性来高效查询目标函数。然而,在许多实际任务中,最高保真函数因成本过高无法参与优化循环。尽管如此,实践者通常拥有先前实验获得的高保真金标准数据。例如在分子优化中,化学家先通过模拟筛选前k个候选分子,再获取其真实目标值。本文证明,在此类现实场景下,即使假设理想,标准MF-BO仍表现不佳。为此,我们提出融合历史高保真数据与任务描述符(可显式提供或从非结构化元数据中提取)的方法,有效缓解此问题。我们在合成函数及化学与超参数优化的真实任务中验证了该方法的有效性。

原文摘要 · Abstract (English)

Black-box optimization is a ubiquitous problem in science and engineering, often dealing with expensive objective functions with cheaper lower-fidelity proxies available. Multi-fidelity Bayesian optimization (MF-BO) is a principled approach to this problem, leveraging correlations across different fidelities when querying the objective. However, for many important MF-BO tasks, the true highest-fidelity function is prohibitively expensive to be part of the optimization loop. Nevertheless, practitioners often have gold standard data (observations of the highest-fidelity function) obtained from previous experiments that might provide information for the current task. For instance, in molecular optimization, chemists often pick the top-$k$ candidate molecules using various computer simulations, and later reveal their true objective function values. In this work, we demonstrate the suboptimality of standard MF-BO algorithms in the real-world scenarios above, even under ideal assumptions. Next, we mitigate this problem by incorporating historical high-fidelity data accompanied by task descriptors---which can be explicitly given or extracted from unstructured metadata. We demonstrate the effectiveness of our methods on synthetic functions, as well as real-world problems in chemistry and hyperparameter optimization.

贝叶斯优化多保真历史数据化学优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。