用可信度动态调整大模型建议,提升多目标优化稳定性
Evidence-Gated LLM Priors for Multi-Objective Bayesian Optimization

- 为每个目标单独建立专家信誉机制,实时更新权重
- 在三个分子优化任务中,动态校准比固定信任更鲁棒
- 大模型信心值不可靠,有时反降低效果,需谨慎使用
大型语言模型(LLMs)被越来越多地用作黑箱优化的启发式顾问,但其建议和自报信心未必与下游目标值一致。这一问题在多目标贝叶斯优化中尤为突出,不同目标可能需要不同专业知识,同一模型对某些目标有用,对另一些则可能误导。本文研究如何在离散多目标贝叶斯优化中使用LLM生成的先验知识,而不盲目信任。提出一种目标级声誉市场机制,将每个专家-目标对视为可验证的先验来源,通过观察到的目标反馈在线更新专家权重,并随时间折扣,同时受市场级信任约束。进一步引入解耦的反事实门控机制,可选择无信心使用、有自信使用或完全不使用LLM先验。在控制性合成压力测试及三个分子优化基准(基于QwenFlash生成先验)上,结果表明动态目标级校准优于固定先验。然而原始的LLM信心值并非始终有益:在ESOL上,信心与预测误差正相关;在FreeSolv上可提升性能;在Lipophilicity上忽略信心表现最佳。固定三臂反事实门控在ESOL和FreeSolv上优于首个变体;而尝试的边际组合暴露了一个重要负结果:边际选择应考虑获取策略,而非仅依赖单步先验误差。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as heuristic advisors for black-box optimization, yet their suggestions and self-reported confidence are not necessarily calibrated to downstream objective values. This issue becomes more pronounced in multi-objective Bayesian optimization, where different objectives may require different expert knowledge and where an LLM expert can be useful for one objective but misleading for another. We study how to use LLM-generated expert priors in discrete multi-objective Bayesian optimization without blindly trusting them. We propose an objective-wise reputation-market mechanism that treats each expert-objective pair as a falsifiable prior source. Expert weights are updated online from observed objective feedback, discounted over time, and gated by market-level trust. We then introduce a decoupled counterfactual gate that can use the LLM prior without confidence, use it with confidence, or abstain from the LLM prior entirely. Across controlled synthetic stress tests and three molecule optimization benchmarks with \qwenflash{}-generated expert priors, we find that dynamic objective-wise calibration improves robustness over fixed LLM priors. However, raw LLM confidence is not reliably beneficial: on ESOL, confidence is positively correlated with prediction error; on FreeSolv, confidence can help; and on Lipophilicity, ignoring confidence remains strongest. Our fixed three-arm counterfactual gate improves over the first counterfactual variant on ESOL and FreeSolv, while an attempted margin portfolio exposes a useful negative result: margin selection should be acquisition-aware rather than based only on one-step prior error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。