arXiv:2510.16165cs.LGcond-mat.supr-con2025-10

评测生成模型重构超导晶体结构的能力,统一比较标准。

AtomBench: A Benchmarking Framework for Generative Crystal Reconstruction Models in Conventional Superconductors

  • 构建统一框架,按推理时信息量分组对比模型
  • MatterGen最准还原原子坐标,FlowMM最快但最不准确
  • 发现临界温度条件对精度无稳定提升,适合材料逆向设计研究者

评估生成式晶体重构模型的关键问题是:提供给模型的晶体学信息量和类型如何影响其重构能力。然而,现有比较常忽略模型在重构过程中接收到的目标信息不均,导致架构结论混淆。我们提出AtomBench,一个可扩展、与模型无关的基准框架,用于在明确定义的晶体重构任务(而非从头生成)上对比生成模型,此处应用于传统超导体。我们在JARVIS Supercon-3D和Alexandria DS-A/B数据集上训练并评估四种模型:AtomGPT、CDVAE、FlowMM和MatterGen,根据它们推理时访问的信息进行分组。重构保真度通过晶格参数的Kullback-Leibler散度(KLD)和平均绝对误差(MAE),以及原子坐标的均方根位移(RMSD)衡量。我们还引入连续校正的RMSD(ccRMSD),为测试集中每个结构定义局部几何保真度的连续度量。结果表明,MatterGen在原子坐标重构上表现最佳,其次为AtomGPT;CDVAE在晶格重构中最为准确,而FlowMM整体最不准确但速度最快。我们发现以临界温度Tc作为条件并未一致提升重构保真度。此外,我们开源了AtomBench Python包,可复现所有报告的重建指标、图表和表格,并支持直接提交至JARVIS-Leaderboard。任何输出晶体重构结果的逆向模型均可使用atombench进行基准测试,欢迎社区使用。

原文摘要 · Abstract (English)

A key question in benchmarking generative crystal reconstruction models is how the amount and type of crystallographic information provided to a generative model affects its ability to reconstruct atomic structures. Yet such comparisons often overlook the fact that models receive unequal information about the target during reconstruction, thereby confounding architectural conclusions. We present AtomBench, an extensible, model-agnostic framework for comparing generative models on a well-defined crystal reconstruction task (rather than \textit{de novo} generation), which we here apply to conventional superconductors. We train and evaluate four models, AtomGPT, CDVAE, FlowMM, and MatterGen, on the JARVIS Supercon-3D and Alexandria DS-A/B datasets, grouping them by the information each accesses at inference. Reconstruction fidelity is measured by the Kullback-Leibler divergence (KLD) and mean absolute error (MAE) of lattice parameters and the root-mean-squared displacement (RMSD) of atomic coordinates. We further introduce the continuous corrected RMSD (ccRMSD), a continuous measure of local geometric fidelity defined for every structure in the test set. MatterGen achieves the best atomic-coordinate reconstruction, followed by AtomGPT, while CDVAE reconstructs lattices most accurately, and FlowMM is the least accurate but fastest overall. We find that conditioning on critical temperature T$_c$ does not consistently improve fidelity. We also release AtomBench as an open-source Python package that reproduces all reported reconstruction metrics, figures, and tables from one or more benchmark files and supports direct submission to the JARVIS-Leaderboard. Any inverse model emitting crystal reconstructions can be benchmarked with \texttt{atombench}, and we encourage community use. https://github.com/atomgptlab/atombench

晶体重构生成模型超导材料基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。