用大模型实现无需训练的材料科学主动学习,实验次数减少70%以上。
Training-Free Active Learning Framework in Materials Science with Large Language Models
- 基于大模型的文本驱动实验设计,无需冷启动和特征工程。
- 在4个数据集上均减少超70%实验量,性能超越传统机器学习方法。
- 适合追求高效、可解释实验设计的研究者,尤其材料领域自动化探索。
主动学习通过优先选择最具信息量的实验加速科学发现,但传统机器学习模型存在冷启动限制和领域特定特征工程问题,制约其泛化能力。大语言模型凭借预训练知识和通用的符号表示,可直接从文本描述中提出实验。本文提出一种基于大模型的主动学习框架(LLM-AL),在少样本迭代设置下运行,并在四个不同材料科学数据集上与传统机器学习模型进行对比。我们测试了两种提示策略:一种采用简洁数值输入,适用于具有更多组成性和结构化特征的数据集;另一种使用扩展性描述文本,适用于具有更多实验和流程特征的数据集以提供上下文。在所有数据集中,LLM-AL将达到最优候选所需的实验次数减少超过70%,并持续优于传统机器学习模型。结果表明,LLM-AL能进行更广泛、更具探索性的搜索,同时以更少迭代次数逼近最优解。我们进一步考察了大模型固有的非确定性对LLM-AL稳定性的影响,发现其性能在多次运行中保持一致,波动范围与传统机器学习方法相当。这些结果表明,LLM-AL可作为通用替代方案,用于更高效、可解释的实验选择,推动大模型驱动的自主发现。
原文摘要 · Abstract (English)
Active learning (AL) accelerates scientific discovery by prioritizing the most informative experiments, but traditional machine learning (ML) models used in AL suffer from cold-start limitations and domain-specific feature engineering, restricting their generalizability. Large language models (LLMs) offer a new paradigm by leveraging their pretrained knowledge and universal token-based representations to propose experiments directly from text-based descriptions. Here, we introduce an LLM-based active learning framework (LLM-AL) that operates in an iterative few-shot setting and benchmark it against conventional ML models across four diverse materials science datasets. We explored two prompting strategies: one using concise numerical inputs suited for datasets with more compositional and structured features, and another using expanded descriptive text suited for datasets with more experimental and procedural features to provide additional context. Across all datasets, LLM-AL could reduce the number of experiments needed to reach top-performing candidates by over 70% and consistently outperformed traditional ML models. We found that LLM-AL performs broader and more exploratory searches while still reaching the optima with fewer iterations. We further examined the stability boundaries of LLM-AL given the inherent non-determinism of LLMs and found its performance to be broadly consistent across runs, within the variability range typically observed for traditional ML approaches. These results demonstrate that LLM-AL can serve as a generalizable alternative to conventional AL pipelines for more efficient and interpretable experiment selection and potential LLM-driven autonomous discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。