arXiv:2509.23156cs.LG2025-09被引 5

用强化学习优化晶体材料,直接反馈密度泛函计算结果。

CrystalGym: A New Benchmark for Materials Discovery Using Reinforcement Learning

  • 构建可交互的强化学习环境,实现晶体设计中DFT反馈闭环。
  • 在多种目标属性上测试算法,发现不同方法样本效率差异显著。
  • 适合关注真实材料设计的机器学习与化学交叉研究者使用。

基于高精度原子模拟器的材料设计主要依赖密度泛函理论(DFT)计算。尽管机器学习可加速材料设计,但多数生成方法因DFT计算成本高,无法直接利用其反馈信号进行训练。为推动在线强化学习中直接使用DFT信号,我们提出CrystalGym——一个开源的晶体材料发现强化学习环境。在该环境中,我们对常见值函数与策略基强化学习算法进行了基准测试,目标是优化带隙、体模量和密度等复杂性质,这些性质均由环境中的DFT直接计算。实验表明,各算法在任务解决率、收敛速度与样本效率方面表现各异。此外,我们还开展大语言模型通过强化学习微调以提升基于DFT奖励性能的案例研究。CrystalGym旨在成为强化学习与材料科学交叉研究的测试平台,为应对高耗时奖励信号的挑战提供新范式,推动面向实际应用的机器学习发展。

原文摘要 · Abstract (English)

In silico design and optimization of new materials primarily relies on high-accuracy atomic simulators that perform density functional theory (DFT) calculations. While recent works showcase the strong potential of machine learning to accelerate the material design process, they mostly consist of generative approaches that do not use direct DFT signals as feedback to improve training and generation mainly due to DFT's high computational cost. To aid the adoption of direct DFT signals in the materials design loop through online reinforcement learning (RL), we propose CrystalGym, an open-source RL environment for crystalline material discovery. Using CrystalGym, we benchmark common value- and policy-based reinforcement learning algorithms for designing various crystals conditioned on target properties. Concretely, we optimize for challenging properties like the band gap, bulk modulus, and density, which are directly calculated from DFT in the environment. While none of the algorithms we benchmark solve all CrystalGym tasks, our extensive experiments and ablations show different sample efficiencies and ease of convergence to optimality for different algorithms and environment settings. Additionally, we include a case study on the scope of fine-tuning large language models with reinforcement learning for improving DFT-based rewards. Our goal is for CrystalGym to serve as a test bed for reinforcement learning researchers and material scientists to address these real-world design problems with practical applications. We therefore introduce a novel class of challenges for reinforcement learning methods dealing with time-consuming reward signals, paving the way for future interdisciplinary research for machine learning motivated by real-world applications.

强化学习材料发现DFT晶体设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。