DASH 解耦优化代理与采集策略,提升自动化贝叶斯优化性能
DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization

- 分离选择代理模型与采集函数,分别基于可靠性与场景上下文优化
- 在四个化学任务中实现12.51%的轨迹加速和5.00%的终点提升
- 结合大模型与知识引导,适合需高效探索的科学实验优化场景
贝叶斯优化依赖代理模型与采集函数,但最佳组合随任务和优化阶段变化。自动化贝叶斯优化(AutoBO)通过在线调整组件应对这一变化,但现有方法或仅调整单一组件,导致另一方不匹配形成瓶颈,或联合选择代理-采集对但忽略其不同作用:代理选择应基于预测可靠性,而采集函数应响应优化上下文。本文提出 DASH,一种用于大语言模型增强型 AutoBO 的解耦自适应代理-采集框架。DASH 通过预测可靠性、不确定性校准和排序一致性选择代理;其两阶段采集控制器定期重新分配各采集函数配额,生成优化候选列表,并交由大语言模型做最终选择。DASH 还集成知识引导冷启动与结构化记忆,融合领域知识与累积反馈。在四个化学优化任务上,DASH 比最优 AutoBO 基线在轨迹级加速因子上提升 12.51%,端点增强因子提升 5.00%。结果在多种大模型底座下保持稳定,消融实验验证各组件互补贡献。全表及行为污染检测未发现直接基准记忆或源细胞泄漏导致性能提升的可检测证据。
原文摘要 · Abstract (English)
Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and optimization stages. Automated Bayesian optimization (AutoBO) addresses this variability by adapting BO components online. However, existing AutoBO methods either adapt one component, leaving the other mismatched and creating a bottleneck, or jointly select surrogate--acquisition pairs under a shared criterion, overlooking their distinct roles: surrogate selection depends on predictive reliability, whereas acquisition adaptation should respond to campaign context.In this paper, we propose DASH, a Decoupled Adaptive Surrogate--Acquisition Harness for large-language- model (LLM)-enhanced AutoBO. DASH selects surrogates by predictive reliability, uncertainty calibration, and ranking consistency; its two-stage acquisition controller periodically reallocates quotas across acquisition functions, builds a BO shortlist accordingly, and delegates final selection to an LLM. DASH also incorporates an integrated harness, consisting of knowledge-guided warm start and structured memory, to ground optimization in domain knowledge and accumulated feedback. Across four chemical optimization tasks, DASH outperforms the best AutoBO baseline by 12.51% in trajectory-level Acceleration Factor and 5.00% in endpoint Enhancement Factor. Results remain strong across LLM backbones, and ablations verify the complementary contributions of all components. Full-table and behavioral contamination checks find no detectable evidence that direct benchmark memorization or source-cell leakage explains these gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。