arXiv:2606.17445cs.LGcond-mat.mtrl-sci2026-06被引 1

用大规模生成模型一键设计满足特定性能的催化剂。

Toward Controllable Catalyst Inverse Design via Large-Scale Autoregressive Pretraining

  • 基于自回归预训练,联合条件生成催化剂结构与性质。
  • 生成结构有效率达98%,吸附能匹配率提升4倍。
  • 适合需要高效筛选反应催化剂的研究者使用。

异相催化剂的逆向设计因表面结构复杂且吸附物相互作用耦合,在广阔的化学空间中难以高效探索。尽管基于机器学习的高通量筛选加速了催化剂发现,但搜索空间扩大时效率仍会下降,促使生成模型直接构建目标性能的催化剂。本文提出一种基于生成预训练变换器架构的条件催化剂生成模型,引入数值嵌入层,在单一自回归框架中同时处理类别与连续属性条件。模型在1.33亿个催化剂结构上预训练,并在约46万条具有类别属性和结合能的优化结构上微调。结果表明,生成结构有效率达98%,优化有效性达95%,类别条件保真度高,吸附物类型与组分联合匹配率达93%;结合能条件匹配率约20%,较基线提升4倍,生成分布系统性逼近目标值,使反应靶向催化剂发现的筛选效率提升1.5至4倍,无需额外微调。这表明大规模自回归预训练结合显式属性条件,为可控催化剂生成与加速发现提供了可行路径。

原文摘要 · Abstract (English)

Inverse design of heterogeneous catalysts remains challenging because catalyst surfaces exhibit substantial structural complexity with coupled surface-adsorbate interactions across a vast chemical space that is difficult to explore efficiently through conventional screening alone. Although machine learning-based high-throughput screening has accelerated catalyst discovery, its efficiency inevitably declines as the search space grows, motivating the development of generative models that can directly construct catalysts with target properties. Here, we present a conditional catalyst generative model based on the Generative Pretrained Transformer architecture with a numerical embedding layer that enables the generation of catalyst structures conditioned on both categorical and continuous properties within a single autoregressive framework. The model was pretrained on 133 million catalyst structures and subsequently fine-tuned on approximately 460,000 optimized structures with associated categorical properties and binding energies for conditional generation. The resulting model achieved 98% structural validity, 95% optimization validity, and high categorical condition fidelity, with a 93 % joint match rate for adsorbate type and composition. For binding energy conditioning, the match rate of approximately 20% represents a four-fold improvement over the baseline training distribution, and the generated distributions shift systematically toward the target values, enabling a 1.5 to 4-fold improvement in screening efficiency for reaction-targeted catalyst discovery without additional fine-tuning. These results show that large-scale autoregressive pre-training, combined with explicit property conditioning, provides a practical route toward controllable catalyst generation and accelerated catalysts discovery.

催化剂设计生成模型逆向设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。