arXiv:2605.26741cond-mat.mtrl-scics.AI2026-05被引 1

首个针对材料逆向设计的系统化评估框架,解决生成算法难比较的问题。

MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation

论文配图:MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation
图 1 · 摘自论文原文
  • 构建物理驱动的合成数据集,模拟真实材料结构-性能关系。
  • 提出多维评估指标,覆盖目标达成率、搜索效率等五个维度。
  • 支持39种算法对比,助力材料逆向设计方法科学验证与优化。

材料逆向设计在目标驱动配方优化方面取得显著进展,但现有机器学习基准仍局限于正向性质预测,未能系统评估逆向优化与生成算法,制约了该领域发展。为此,我们提出MatFormBench,一个专为评估和指导目标驱动配方生成策略而设计的新型基准生态体系。该框架融合物理驱动的配方生成机制,生成忠实模拟真实材料结构-性能响应关系的合成样本,并设置五级递增难度以量化关系复杂性。为严格评估算法性能,我们进一步提出MatFormScore,一个涵盖目标达成率、搜索效率、探索能力、鲁棒性和稳定性五个关键维度的多维评价指标。通过评估39种不同逆向设计算法(包括经典代理辅助黑箱搜索、先进深度生成模型及流行的大型语言模型推荐策略),在1170项标准化算法-任务评估中,基于扩散模型的方法表现最优;变分自编码器(VAE)与遗传算法(GA)则在特定场景下具优势。MatFormBench建立统一评估标准,实现可复现的基准测试、合理算法比较与诊断分析,为推动材料逆向设计提供基础工具。

原文摘要 · Abstract (English)

Inverse design of materials has significantly advanced target-driven formulation optimization, yet existing materials machine learning benchmarks remain limited to forward property prediction, failing to systematically evaluate inverse optimization and generation algorithms, a critical gap that hinders the progress of target-driven materials design. To address this limitation, we propose MatFormBench, a novel benchmarking ecosystem tailored to evaluate and guide generative strategies for target-driven formulation. MatFormBench integrates a physics-driven formulation generation scheme to generate synthetic samples that faithfully emulate realistic materials structure-property response relationships, complemented by five escalating difficulty levels to quantify the complexity of these relationships. To rigorously assess algorithm performance, we further propose MatFormScore, a multi-dimensional metric that comprehensively quantifies performance across five critical axes: target success, search efficiency, exploratory capacity, robustness, and stability. We validate MatFormBench by evaluating 39 diverse inverse design algorithms, covering classical surrogate-assisted black-box search, state-of-the-art deep generative models, and increasingly popular Large Language Model (LLM)-based recommendation strategies. Across 1170 standardized algorithm-task evaluations, diffusion-based models demonstrate the strongest overall performance, while Variational Autoencoder (VAE)-based and Genetic Algorithm (GA)-based methods exhibit distinct advantages in specific scenarios. By establishing a unified evaluation standard for target-driven materials formulation, MatFormBench enables reproducible benchmarking, principled algorithm comparison, and diagnostic analysis of inverse design strategies, providing a foundational tool for advancing materials inverse design.

材料逆向设计生成模型评估基准扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。