arXiv:2603.12567cond-mat.mtrl-scics.LG2026-03

用预训练大模型做材料发现的智能筛选,大幅减少实验次数。

Foundation-Model Surrogates Enable Data-Efficient Active Learning for Materials Discovery

  • 用基于变压器的预训练模型替代传统预测器,无需重新训练
  • 在10个数据集上平均节省52%实验次数,优于传统方法
  • 适合小样本实验场景,尤其适用于数据稀缺的材料研究

主动学习(AL)通过迭代引导实验寻找优质材料候选,显著减少昂贵的合成与表征循环次数。然而,当前方法主要依赖高斯过程(GP)和随机森林(RF)作为代理模型,存在明显局限:GP因核函数假设僵化,在复杂成分-性能关系中表现不佳;而RF在小样本下不确定性估计不可靠。这一小样本挑战在材料科学中普遍存在,导致从零开始训练模型难以保证可靠性。本文提出上下文主动学习(ICAL),以TabPFN——一个在百万个合成回归任务上预训练的变压器基础模型——取代传统代理模型。该模型能通过单次前向传播实现严谨的贝叶斯推断,无需针对特定数据集再训练,兼具优异的小样本回归性能和校准良好的预测不确定性(主动学习所需)。我们在10个材料数据集上对比了ICAL、GP和RF,TabPFN在8个数据集上表现更优,相比GP平均减少52%额外评估,相比RF减少29.77%。交叉验证分析表明,优势源于更优的不确定性校准,其负对数似然和稀疏误差曲线下面积均最低。结果证明,预训练基础模型可作为高效主动学习代理,推动多样材料体系及小样本实验科学中的数据高效发现。

原文摘要 · Abstract (English)

Active learning (AL) has emerged as a powerful paradigm for accelerating materials discovery by iteratively steering experiments toward promising candidates, reducing the number of costly synthesis-and-characterization cycles needed to identify optimal materials. However, current AL relies predominantly on Gaussian Process (GP) and Random Forest (RF) surrogates, which suffer from complementary limitations: GP underfits complex composition-property landscapes due to rigid kernel assumptions, while RF produces unreliable heuristic uncertainty estimates in small-data regimes. This small-data challenge is pervasive in materials science, making reliable surrogate modeling extremely difficult with models trained from scratch on each new dataset. Here we propose In-Context Active Learning (ICAL), which addresses this bottleneck by replacing conventional surrogates with TabPFN, a transformer-based foundation model (FM) pre-trained on millions of synthetic regression tasks to meta-learn a universal prior over tabular data, upon which TabPFN performs principled Bayesian inference in a single forward pass without dataset-specific retraining, delivering strong small-data regression performance and well-calibrated predictive uncertainty (required for effective AL). We benchmark ICAL against GP and RF across 10 materials datasets and TabPFN wins on 8 out of 10 datasets, achieving a mean saving of 52% in extra evaluations relative to GP and 29.77% relative to RF. Cross-validation analysis confirms that TabPFN's advantage stems from superior uncertainty calibration, achieving the lowest Negative Log-Likelihood and Area Under the Sparsification Error curve among all surrogates. These results demonstrate that pre-trained FMs can serve as effective surrogates for active learning, enabling data-efficient discovery across diverse materials systems and small-data experimental sciences.

主动学习材料发现小样本预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。