构建闭环材料发现基准,评估自动化流程效率与适应性。
MADE: Benchmark Environments for Closed-Loop Materials Discovery
- 模拟资源受限的迭代式材料发现过程,定义为稳定相搜索
- 通过对比基线算法验证不同方法在复杂度下的性能差异
- 支持模块化组件组合,适合研究智能代理与自适应决策
现有材料发现基准多聚焦静态预测任务或孤立计算子任务,忽视科学发现的迭代与适应特性。我们提出MAterials Discovery Environments(MADE),一个用于评估端到端自主材料发现流程的新框架。MADE模拟在有限查询预算下,智能体提出、评估并优化候选材料的闭环发现过程,捕捉真实科研工作中的序列性与资源约束。将发现形式化为相对于给定凸包的热力学稳定化合物搜索,并通过与基线算法比较来评估有效性与效率。该框架具有高度灵活性,用户可自由组合生成模型、过滤器与规划器等模块,研究从固定流程到具备工具使用与自适应决策能力的全智能代理系统。我们通过系统实验在多个体系上展开,实现对发现流程组件的消融分析,并考察方法随系统复杂度的扩展表现。
原文摘要 · Abstract (English)
Existing benchmarks for computational materials discovery primarily evaluate static predictive tasks or isolated computational sub-tasks. While valuable, these evaluations neglect the inherently iterative and adaptive nature of scientific discovery. We introduce MAterials Discovery Environments (MADE), a novel framework for benchmarking end-to-end autonomous materials discovery pipelines. MADE simulates closed-loop discovery campaigns in which an agent or algorithm proposes, evaluates, and refines candidate materials under a constrained oracle budget, capturing the sequential and resource-limited nature of real discovery workflows. We formalize discovery as a search for thermodynamically stable compounds relative to a given convex hull, and evaluate efficacy and efficiency via comparison to baseline algorithms. The framework is flexible; users can compose discovery agents from interchangeable components such as generative models, filters, and planners, enabling the study of arbitrary workflows ranging from fixed pipelines to fully agentic systems with tool use and adaptive decision making. We demonstrate this by conducting systematic experiments across a family of systems, enabling ablation of components in discovery pipelines, and comparison of how methods scale with system complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。