用强化学习提升晶体结构生成质量,兼顾准确率与多样性。
CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure Prediction

- 联合坐标与晶格状态的强化学习策略,优化生成过程
- 在四个数据集上均降低1样本和20样本的均方误差
- 适合需要高精度和多样性的晶体结构预测任务
基于流的生成模型可高效生成晶体结构候选,但其预训练目标不直接优化下游目标结构的恢复。强化学习后训练提供灵活解决方案,但现有方法主要依赖能量奖励和仅坐标随机策略。预测能量无法识别参考多晶型,而奖励驱动集中会降低实现Top-N恢复所需的候选覆盖率。本文提出CrystalGRPO,一种面向晶体结构预测的后训练框架,将现有从常微分方程到随机微分方程的策略构造扩展至坐标-晶格联合状态。CrystalGRPO结合MACE预测能量与StructureMatcher恢复评分,提供两种模式:CrystalGRPO-Q优先单次采样恢复,CrystalGRPO-C则结合完整轨迹参考正则化与覆盖感知组优势,以保持有限预算下的目标恢复覆盖率。在MP-20与MPTS-52数据集上,采用PXRDGen与OMatG骨干网络,两种变体在所有四组骨干-数据集设置中均降低了1样本与20样本的均方误差。CrystalGRPO-Q始终提升Top-1表现,CrystalGRPO-C在所有设置中均获得更高Top-20性能。
原文摘要 · Abstract (English)
Flow-based generative models can efficiently produce candidate structures for crystal structure prediction (CSP), but their pretrained objectives do not directly optimize downstream target recovery. Reinforcement-learning post-training offers a flexible solution, yet existing approaches rely primarily on energy rewards and coordinate-only stochastic policies. Predicted energy does not identify the reference polymorph, while reward-driven concentration can reduce the candidate coverage required for Top-N recovery. We introduce CrystalGRPO, a CSP-aligned post-training framework that extends existing ODE-to-SDE policy constructions to the joint coordinate--lattice state. CrystalGRPO combines MACE-predicted energy with a StructureMatcher-based recovery score and provides two operating modes: CrystalGRPO-Q, which prioritizes single-draw recovery, and CrystalGRPO-C, which combines full-trajectory reference regularization with a coverage-aware group advantage to preserve finite-budget target recovery. Across MP-20 and MPTS-52 with PXRDGen and OMatG backbones, both variants reduce one- and twenty-sample RMSE relative to coordinate-only reinforcement in all four backbone--dataset settings. CrystalGRPO-Q consistently improves Top-1, whereas CrystalGRPO-C achieves a higher Top-20 across all settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。