arXiv:2504.02367cond-mat.mtrl-scics.LG2025-04被引 17

用强化学习优化材料生成模型,同时实现高介电常数与宽能隙的难兼得材料设计。

Reinforcement Fine-Tuning for Materials Design

  • 以判别模型奖励信号指导生成模型,融合多属性优化目标。
  • 生成晶体稳定性提升,成功发现兼具高介电常数和宽能隙的材料。
  • 让预训练模型具备按性质检索材料的能力,拓展应用场景。

强化微调在提升大语言模型指令遵循与推理能力方面发挥了关键作用。本文将强化微调应用于材料设计,利用判别式机器学习模型为基于自回归Transformer的材料生成模型CrystalFormer提供奖励信号。通过优化能量凸包以上值及材料性能指标等奖励目标,强化微调将判别模型的知识注入生成模型。所得模型CrystalFormer-RL在生成晶体时表现出更高稳定性,并成功发现具有理想但相互冲突的材料性质(如同时具备高介电常数和宽带隙)的晶体。值得注意的是,强化微调不仅实现了属性导向的材料设计,还激发了预训练生成模型的属性驱动材料检索行为。该框架为机器学习生态在材料设计中的协同应用开辟了新路径。

原文摘要 · Abstract (English)

Reinforcement fine-tuning played an instrumental role in enhancing the instruction-following and reasoning abilities of large language models. In this work, we employ reinforcement fine-tuning for materials design, in which discriminative machine learning models are used to provide rewards to the autoregressive transformer-based materials generative model CrystalFormer. By optimizing the reward signals-such as energy above the convex hull and material properties figures of merit-reinforcement fine-tuning infuses knowledge from discriminative models into generative models. The resulting model, CrystalFormer-RL, shows enhanced stability in generated crystals and successfully discovers crystals with desirable yet conflicting material properties, such as substantial dielectric constant and band gap simultaneously. Notably, we observe that reinforcement fine-tuning not only enables the property-guided material design but also unlocks property-based material retrieval behavior of pretrained generative model. The present framework opens an exciting gateway to the synergies of the machine learning ecosystem for materials design.

材料生成强化学习晶体设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。