建立统一评估框架,让晶体生成模型比拼更公平可靠。
LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models
- 提出多维度评估指标,统一衡量生成模型性能。
- 12个模型测试显示稳定性提升会降低新颖性和多样性。
- 开源工具与排行榜助力科研界持续优化生成模型。
生成式机器学习模型在通过逆向设计无机晶体加速材料发现方面具有巨大潜力,能前所未有地探索化学空间。然而,缺乏标准化的评估框架使得模型评价、比较和进一步发展难以实现。本文提出LeMat-GenBench,一个面向晶体生成模型的统一基准,配套设计了多项评估指标,以更好指导模型开发和下游应用。我们开源了评估工具包,并在Hugging Face上发布公开排行榜,对12个近期生成模型进行了基准测试。结果表明,平均而言,模型稳定性提升会导致新颖性和多样性下降,且无一模型在所有维度上均表现优异。总体而言,LeMat-GenBench为可复现、可扩展的公平模型比较奠定了基础,旨在引导开发更可靠、以发现为导向的晶体生成模型。
原文摘要 · Abstract (English)
Generative machine learning (ML) models hold great promise for accelerating materials discovery through the inverse design of inorganic crystals, enabling an unprecedented exploration of chemical space. Yet, the lack of standardized evaluation frameworks makes it challenging to evaluate, compare, and further develop these ML models meaningfully. In this work, we introduce LeMat-GenBench, a unified benchmark for generative models of crystalline materials, supported by a set of evaluation metrics designed to better inform model development and downstream applications. We release both an open-source evaluation suite and a public leaderboard on Hugging Face, and benchmark 12 recent generative models. Results reveal that an increase in stability leads to a decrease in novelty and diversity on average, with no model excelling across all dimensions. Altogether, LeMat-GenBench establishes a reproducible and extensible foundation for fair model comparison and aims to guide the development of more reliable, discovery-oriented generative models for crystalline materials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。