arXiv:2509.08736cs.LG2025-09被引 3

用大模型多智能体系统加速化学贝叶斯优化,效率最高提升5倍。

ChemBOMAS: Accelerated BO in Chemistry with LLM-Enhanced Multi-Agent System

  • 引入80亿参数大模型生成伪数据,仅用1%标注样本即可初始化优化。
  • 结合检索增强生成与置信上界算法,将搜索空间分块并定位高潜力区域。
  • 在多个科学基准上实现最优性能,适合高效化学实验设计的研究者。

贝叶斯优化(BO)是化学科学发现的强大工具,但其效率常受实验数据稀疏和搜索空间庞大限制。本文提出ChemBOMAS:一种大语言模型(LLM)增强的多智能体系统,通过数据与知识协同驱动策略加速BO。首先,数据驱动策略采用在仅1%标注样本上微调的80亿参数LLM回归器生成伪数据,有效初始化优化过程;其次,知识驱动策略利用混合检索增强生成方法引导LLM划分搜索空间,并抑制幻觉。随后,置信上界(UCB)算法在已划分的子空间中识别高潜力区域。在经LLM优化的子空间及生成数据支持下,BO显著提升效果与效率。多组科学基准评估表明,ChemBOMAS达到新最优水平,相比基线方法最多提升5倍优化效率。

原文摘要 · Abstract (English)

Bayesian optimization (BO) is a powerful tool for scientific discovery in chemistry, yet its efficiency is often hampered by the sparse experimental data and vast search space. Here, we introduce ChemBOMAS: a large language model (LLM)-enhanced multi-agent system that accelerates BO through synergistic data- and knowledge-driven strategies. Firstly, the data-driven strategy involves an 8B-scale LLM regressor fine-tuned on a mere 1% labeled samples for pseudo-data generation, robustly initializing the optimization process. Secondly, the knowledge-driven strategy employs a hybrid Retrieval-Augmented Generation approach to guide LLM in dividing the search space while mitigating LLM hallucinations. An Upper Confidence Bound algorithm then identifies high-potential subspaces within this established partition. Across the LLM-refined subspaces and supported by LLM-generated data, BO achieves the improvement of effectiveness and efficiency. Comprehensive evaluations across multiple scientific benchmarks demonstrate that ChemBOMAS set a new state-of-the-art, accelerating optimization efficiency by up to 5-fold compared to baseline methods.

贝叶斯优化大模型化学AI多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。