arXiv:2605.30889physics.chem-phcs.LG2026-05被引 1

用大模型自动优化原子间势能模型,让代码自动生成、实验自动执行。

MLIPilot: LLM-Driven Auto-Research for Machine-Learned Interatomic Potentials

论文配图:MLIPilot: LLM-Driven Auto-Research for Machine-Learned Interatomic Potentials
图 1 · 摘自论文原文
  • 大模型调用工具改代码、跑计算,按物理约束评分
  • 在两种数据集上成功修复初始不稳定的基线模型
  • 适合想自动化材料模拟研发的科研人员

构建高质量机器学习原子间势能模型需在精度、动力学稳定性和计算效率之间权衡,而这些约束无法通过单一损失函数体现。我们提出 MLIPilot,一个由工具调用的大语言模型驱动的自动研究框架:模型提出假设、修改训练代码、提交高性能计算任务,并根据固定的物理约束评分卡决定是否保留或回滚更改。我们在 MACE 模型优化中评估了包括 GPT-5.5、GPT-4.1、Mistral-24B 和 Qwen3-32B 在内的商业与开源大模型代理。测试涵盖分子与周期性体系:基于 QM7 的数据集(使用 B3LYP/6-31G(d) 计算能量与力),以及由 ASE 的有效介质理论计算器标注的周期性铜超胞数据集。在这些基准上,最强代理通过发现输出归一化、损失函数调整、渐进式训练策略和模型容量优化等有效方法,将初始违反约束的基线模型成功改进为可接受模型。结果表明,在领域特定验证标准约束下,大模型可作为科学机器学习流程的自主操作员,将部分 MLIP 开发从手动试错转向可审计的自动化实验。

原文摘要 · Abstract (English)

Constructing production-quality machine-learned interatomic potentials (MLIPs) requires balancing accuracy, dynamical stability, and computational throughput under constraints that are not captured by a single training loss. We introduce MLIPilot, an auto-research framework in which tool-calling large language models propose hypotheses, edit MLIP training code, launch HPC jobs, and accept or revert changes using a fixed, physically constrained scorecard. We evaluate MLIPilot on MACE potential optimization using both commercial and open-weight LLM agents, including GPT-5.5, GPT-4.1, Mistral-24B, and Qwen3-32B. The benchmarks span molecular and periodic settings: a QM7-derived dataset for which we generated B3LYP/6-31G(d) energies and forces, and a Cu EMT dataset with periodic copper supercells labeled by ASE's Effective Medium Theory calculator. Across these benchmarks, the strongest agents move initially constraint-violating baselines to accepted models by discovering useful training strategies, including output normalization, loss-function changes, progressive training schedules, and model-capacity adjustments. These results suggest that LLM agents can serve as autonomous operators for scientific machine-learning workflows when their search is constrained by domain-specific validation criteria, shifting part of MLIP development from manual trial-and-error toward auditable, automated experimentation.

机器学习势能大模型编程自动化科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。