arXiv:2507.14267cs.AIcond-mat.mtrl-sci2025-07被引 35

DREAMS让材料模拟机器人可靠执行复杂计算,每步都可追溯、可验证。

DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation

  • 构建多层级安全守卫,逐项检查数值来源与逻辑
  • 在晶格常数测试中误差低于1%,吸附能预测达专家水平
  • 适合需要高可信度的自动化材料研发团队使用

大语言模型代理可执行长周期科学工作流,但其数值结果难以信任:易丢失上下文、绕过验证且生成大量看似合理实则错误的结果。我们提出基于密度泛函理论(DFT)的智能材料模拟研究引擎DREAMS,一个分层多代理框架,围绕多级安全守卫构建。守卫在存在明确标准处实施确定性检查,在其他区域采用受限的LLM判断,逐参数评估并追踪每个值的注册来源。验证覆盖从工具调用时(拒绝伪造、篡改或无源数据)到报告生成时(审计每项结论背后的完整溯源图),共享画布保障数百步操作中的信息完整性。DREAMS在Sol27LC晶格常数基准上平均误差低于1%,复现了专家级的CO/Pt(111)吸附能差异,并通过贝叶斯集成采样量化泛函驱动的不确定性,证实广义梯度近似(GGA)水平下面心立方(FCC)位点偏好。相比未加防护的对照系统(仅81%关键步骤成功却得近似正确答案),该系统以约13倍输入令牌数完成全程验证;验证层级可独立关闭,实现可信度与成本的权衡,调优后的判断规则可在五个不同判别模型间迁移。DREAMS达到增强型L2(L2+)自动化水平,展现出接近L3自动化的潜力,为可信、高通量自主材料模拟提供可行路径。

原文摘要 · Abstract (English)

Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose context, game verification checks, and can produce large volumes of plausible yet invalid results. We introduce the DFT-based Research Engine for Agentic Materials Simulation (DREAMS), a hierarchical multi-agent framework for density functional theory (DFT) built around a multi-tier safety guard. The guard applies deterministic checks wherever explicit criteria exist and scoped LLM judgment elsewhere, evaluating one parameter at a time and tracing every value to its registered source. Verification extends from tool-call time, where fabricated, laundered, or unsourced values are rejected before entering the workflow, to report time, where a judge audits the full provenance graph behind every claim; a shared canvas preserves information integrity across hundreds of steps. DREAMS achieves average errors below 1% on the Sol27LC lattice-constant benchmark, reproduces expert-level adsorption-energy differences on the CO/Pt(111) puzzle, and quantifies functional-driven uncertainty with Bayesian ensemble sampling, confirming the face-centered-cubic (FCC) site preference at the generalized gradient approximation (GGA) level. Compared with its unguarded counterpart, which reached a nearly correct answer while only 81% of its essential steps succeeded, the guarded system verifies every essential step at approximately 13 times the input tokens; verification layers can be disabled individually to balance trustworthiness against cost, and the tuned judge rules transfer across five judge models. DREAMS operates at an enhanced L2 (L2+) automation level and demonstrates capabilities approaching L3 automation, providing a path toward trustworthy, high-throughput autonomous materials simulation.

材料模拟AI代理可靠性验证DFT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。