用大模型做材料优化选择,无需训练也能有效找最优解。
LLMs as Acquisition Policies for Finite-Pool Materials Optimization: A Controlled Study

- 直接用大模型决定下一步实验候选,不需任务特定训练。
- 相比随机选择,更快找到全局最优,但表现因任务而异。
- 适合材料科学中低成本试错场景,尤其当数据有限时。
发现具有理想性能的材料通常需要在庞大候选空间中搜索,而实验或计算评估成本高昂。主动学习通过利用先前观测结果来选择下一步评估的候选者,通常依赖概率代理模型。本文研究开放权重的大语言模型(LLMs)是否可作为此类场景下的独立获取策略。我们在四种回溯性有限池材料优化任务上评估了五种LLM,采用不同候选呈现策略,并与随机选择和传统的高斯过程方法进行对比。结果显示,LLM策略在多数情况下比随机选择更早达到全局最优,表明其能提供有效的获取信号且无需任务特异性训练。与高斯过程方法相比,表现参差不齐:传统获取策略在大多数任务上更优,但在某些设置下LLM表现相当甚至更好。性能在任务、模型、初始化及候选呈现方式间差异显著,无单一LLM在所有任务中均最优。总体而言,开放权重的LLM在有限池材料搜索中具备潜力,但其可靠性仍受任务特性及候选与科学背景呈现方式的影响。
原文摘要 · Abstract (English)
Discovering materials with desirable properties often requires searching large candidate spaces while experimental or computational evaluations remain costly. Active learning addresses this challenge by using previous observations to select which candidate to evaluate next, typically through probabilistic surrogate models. We investigate whether open-weight large language models (LLMs) can serve as standalone acquisition policies in this setting. We evaluate five LLMs across four retrospective finite-pool materials optimization tasks under different candidate-presentation strategies and compare them with random selection and conventional Gaussian-process methods. LLM policies generally reach the global optimum in fewer iterations than random selection, indicating that they provide a useful acquisition signal without task-specific training. Their performance relative to Gaussian-process methods is mixed: conventional acquisition performs better on most tasks, while LLMs match or outperform it in some settings. Performance varies substantially across tasks, models, initializations, and candidate presentations, with no LLM approach performing best across all tasks. Overall, open-weight LLMs show potential as acquisition policies for finite-pool materials search, although their reliability remains sensitive to the task and to how candidates and scientific context are presented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。