arXiv:2510.18900physics.chem-phcond-mat.mtrl-sci2025-10被引 10

MIST模型通过大规模分子预训练,实现化学空间高效探索与多任务预测。

Foundation Models for Discovery and Exploration in Chemical Space

  • 基于新型分词器Smirk,融合核、电子与几何信息构建分子表示。
  • 在400+结构-性能关系上表现超主流方法,涵盖电化学到生理学领域。
  • 可解决未预设的复杂问题,如气味感知建模与立体化学推理。

从分子结构准确预测原子、热力学和动力学性质是材料创新的基础。现有计算与实验方法难以高效探索化学空间。基于大规模无标签数据训练的科学基础模型为跨领域导航化学空间提供了可能。本文提出MIST系列分子基础模型,参数量和数据规模较此前工作提升一个数量级。采用新型分词器Smirk,全面捕捉核、电子与几何信息,使模型学习多样化分子表征。经微调后,MIST可预测超过400种结构-性能关系,在从生理学到电化学的多个基准测试中达到或超越当前最优水平。我们展示了其在真实问题中的应用能力,包括多目标电解质溶剂筛选、有机金属化合物立体化学推理及混合物性质预测。最显著体现基础模型潜力的是解决非训练目标的问题——我们识别出嗅觉感知映射为此类问题,发现MIST能准确预测气味特征,并学习到符合双曲几何的嗅觉空间层次结构。此外,我们提出了超参数感知的贝叶斯神经网络缩放定律,避免在每尺度下进行超参搜索,使在有限算力下训练高计算效率的大模型成为可能。本研究为利用基础模型加速材料发现、设计与优化迈出重要一步。

原文摘要 · Abstract (English)

Accurate prediction of atomistic, thermodynamic, and kinetic properties from molecular structures underpins materials innovation. Existing computational and experimental approaches lack the scalability required to navigate chemical space efficiently. Scientific foundation models trained on large unlabelled datasets offer a path towards navigating chemical space across application domains. Here, we develop MIST, a family of molecular foundation models with up to an order of magnitude more parameters and data than prior works. Trained using a novel tokenizer, Smirk, which comprehensively captures nuclear, electronic, and geometric information, MIST learns a diverse range of molecules. MIST models have been fine-tuned to predict more than 400 structure-property relationships and have been shown to match or exceed state-of-the-art performance across diverse benchmarks, from physiology to electrochemistry. We demonstrate the ability of these models to solve real-world problems across chemical space from multiobjective electrolyte solvent screening to stereochemical reasoning for organometallics and mixture property prediction. The clearest demonstration of a foundation model is its ability to solve problems that were neither explicit targets of training nor central to the intentions of its developers. We identify olfactory perception mapping as such a problem, and show that MIST accurately predicted scent profiles and learned a hierarchical representation of olfactory space consistent with hyperbolic geometry. We formulated hyperparameter aware Bayesian neural scaling laws which eliminate the need for hyperparameter sweeps at every scale, making training large compute-optimal models feasible on a limited compute budget. The methods and findings presented here represent a significant step towards accelerating materials discovery, design, and optimization using foundation models.

分子建模基础模型化学空间多任务预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。