arXiv:2509.00103cs.LGcs.AI2025-09被引 3

大模型预训练知识让化学反应优化突破传统算法瓶颈。

Pre-trained knowledge elevates large language models beyond traditional chemical reaction optimizers

  • 用大模型引导搜索,利用领域知识高效探索化学参数空间。
  • 在复杂高维分类空间中,大模型性能超越贝叶斯优化,尤其在优质解稀少时优势显著。
  • 开源平台支持透明对比,适合化学研发与算法优化研究者使用。

现代实验化学中的优化依赖于黑箱参数空间的算法搜索。本文展示大语言模型(LLM)中的预训练知识从根本上改变了这一范式。基于六个完整枚举的分类反应数据集(768-5,684次实验),我们对比了大模型引导优化(LLM-GO)与贝叶斯优化(BO)及随机采样。前沿大模型在五个单目标数据集上表现持平或优于BO,且随着参数复杂度增加、优质条件稀少(<5%空间)时优势更明显;仅在显式多目标权衡场景中,BO仍占优。为理解差异,我们提出一种拓扑无关的信息论框架,量化优化全过程的采样多样性。分析表明,所有数据集中,大模型维持更高的探索香农熵,同时实现更优性能,尤其在高熵探索通常失效的稀缺解空间中——说明预训练知识提升了对化学空间的导航能力,而非替代结构化探索策略。为促进透明评估与社区验证,我们发布Iron Mind(https://gomes.andrew.cmu.edu/iron-mind)平台,支持人、算法与大模型优化方案的并行评估,含公开排行榜与完整轨迹。研究证实,LLM-GO在传统方法难以应对的复杂分类空间中表现卓越:需领域理解而非数学优化的场景。

原文摘要 · Abstract (English)

Modern optimization in experimental chemistry employs algorithmic search through black-box parameter spaces. Here we demonstrate that pre-trained knowledge in large language models (LLMs) fundamentally changes this paradigm. Using six fully enumerated categorical reaction datasets (768-5,684 experiments), we benchmark LLM-guided optimization (LLM-GO) against Bayesian optimization (BO) and random sampling. Frontier LLMs consistently match or exceed BO performance across five single-objective datasets, with advantages growing as parameter complexity increases and high-performing conditions become scarce (<5% of space). BO retains superiority only for explicit multi-objective trade-offs. To understand these contrasting behaviors, we introduce a topology-agnostic information theory framework quantifying sampling diversity throughout optimization campaigns. This analysis reveals that LLMs maintain systematically higher exploration Shannon entropy than BO across all datasets while achieving superior performance, with advantages most pronounced in solution-scarce parameter spaces where high-entropy exploration typically fails-suggesting that pre-trained domain knowledge enables more effective navigation of chemical parameter space rather than replacing structured exploration strategies. To enable transparent benchmarking and community validation, we release Iron Mind (https://gomes.andrew.cmu.edu/iron-mind), a no-code platform for side-by-side evaluation of human, algorithmic, and LLM optimization campaigns with public leaderboards and complete trajectories. Our findings establish that LLM-GO excels precisely where traditional methods struggle: complex categorical spaces requiring domain understanding rather than mathematical optimization.

大模型化学优化预训练知识智能实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。