AI自主发现材料科学理论,能自动推导方程并验证。
From Data to Theory: Autonomous Large Language Model Agents for Materials Science
- 用大模型自主选方程、写代码、验数据,全程无需人工干预。
- 对霍尔-佩奇等经典公式能准确复现,对新关系可提出假设。
- 适合科研人员探索未知规律,但需人工验证结果可靠性。
我们提出一种自主的大型语言模型(LLM)代理,用于端到端、数据驱动的材料科学理论构建。该模型可自主选择方程形式,生成并运行代码,无需人工干预即可测试理论与数据的匹配程度。框架结合逐步推理与专家提供的工具,使代理能动态调整策略,并保留决策记录。对于霍尔-佩奇方程、巴黎定律等成熟材料关系,代理能正确识别控制方程并在新数据集上做出可靠预测。对于更专门的关系,如共轭分子长度与HOMO-LUMO间隙的库恩方程,性能更依赖底层模型,其中GPT-5表现更好。此外,代理还能提出新的预测关系,例如应变依赖的HOMO-LUMO间隙变化规律。然而,结果也表明,严谨验证仍至关重要,因为即使数值拟合良好,代理也可能返回错误、不完整或不一致的方程。总体而言,这些结果揭示了自主LLM代理在人工智能辅助科学建模与发现中的潜力与当前局限。
原文摘要 · Abstract (English)
We present an autonomous large language model (LLM) agent for end-to-end, data-driven materials theory development. The model can choose an equation form, generate and run its own code, and test how well the theory matches the data without human intervention. The framework combines step-by-step reasoning with expert-supplied tools, allowing the agent to adjust its approach as needed while keeping a clear record of its decisions. For well-established materials relationships such as the Hall-Petch equation and Paris law, the agent correctly identifies the governing equation and makes reliable predictions on new datasets. For more specialized relationships, such as Kuhn's equation for the HOMO-LUMO gap of conjugated molecules as a function of length, performance depends more strongly on the underlying model, with GPT-5 showing better recovery of the correct equation. Beyond known theories, the agent can also suggest new predictive relationships, illustrated here by a strain-dependent law for changes in the HOMO-LUMO gap. At the same time, the results show that careful validation remains essential, because the agent can still return incorrect, incomplete, or inconsistent equations even when the numerical fit appears strong. Overall, these results highlight both the promise and the current limitations of autonomous LLM agents for AI-assisted scientific modeling and discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。