用大模型理解自然语言反馈,让优化更贴近人类主观判断。
LILO: Bayesian Optimization with Natural Language Feedback
- 用大模型将自然语言反馈转化为结构化偏好信号
- 在有限反馈下表现优于传统方法和纯大模型优化
- 适合需要主观判断的工程优化场景
许多现实优化问题受复杂主观偏好驱动,难以用明确数学目标表达。为此,我们提出语言在环优化(LILO),一种贝叶斯优化框架,通过大语言模型(LLM)将决策者提供的自由形式自然语言反馈与先验知识转化为结构化偏好信号,突破了传统偏好贝叶斯优化中仅支持标量或成对反馈的限制。由LLM生成的偏好通过高斯过程代理模型整合,实现带有校准不确定性的合理采集驱动探索。通过将LLM作为辅助而非主优化器,LILO保持了贝叶斯优化的样本效率与稳定性,同时提供灵活且表达力强的反馈接口。在合成与真实世界基准测试中,LILO始终优于传统偏好型贝叶斯优化方法及纯大模型优化器,尤其在反馈稀缺场景下提升显著。
原文摘要 · Abstract (English)
Many real-world optimization problems are guided by complex, subjective preferences that are difficult to express as explicit closed-form objectives. In response, we introduce Language-in-the-Loop Optimization (LILO), a Bayesian optimization (BO) framework that employs a large language model (LLM) to translate free-form natural language feedback and prior knowledge from a decision maker into structured preference signals, going beyond the restrictive scalar or pairwise feedback formats typically assumed in preferential BO. The LLM-derived preferences are integrated by a Gaussian process proxy model, enabling principled acquisition-driven exploration with calibrated uncertainty. By placing the LLM in a supporting role rather than as the optimizer itself, LILO preserves the sample efficiency and stability of BO while providing a flexible and expressive feedback interface. Across synthetic and real-world benchmarks, LILO consistently outperforms both conventional preference-based BO methods and LLM-only optimizers, with particularly strong gains in feedback-limited regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。