用大模型自动搭建蛋白分子动力学模拟,省时又少错。
Automating MD simulations for Proteins using Large language Models: NAMD-Agent
- 用Gemini大模型+脚本自动操作CHARMM GUI生成模拟文件。
- 相比人工操作,设置时间显著缩短,错误率大幅降低。
- 适合需要批量处理蛋白系统的计算生物学家使用。
分子动力学模拟是理解蛋白质原子层面结构、动态与功能的重要工具,但准备高质量的输入文件耗时且易出错。本文提出一种自动化流程,利用大语言模型Gemini 2.0 Flash结合Python脚本与Selenium网页自动化技术,通过CHARMM GUI的网页界面自动生成适用于NAMD的模拟输入文件。该流程借助Gemini的代码生成与迭代优化能力,自动编写、执行并修正模拟脚本,完成对CHARMM GUI的操作、参数提取及输入文件生成。后续通过其他软件进行后处理,实现从准备到输出的全流程自动化。结果表明,该方法显著减少设置时间,降低人为错误,支持多蛋白系统并行处理,为计算结构生物学中大模型应用提供可扩展、可靠的自动化平台。
原文摘要 · Abstract (English)
Molecular dynamics simulations are an essential tool in understanding protein structure, dynamics, and function at the atomic level. However, preparing high quality input files for MD simulations can be a time consuming and error prone process. In this work, we introduce an automated pipeline that leverages Large Language Models (LLMs), specifically Gemini 2.0 Flash, in conjunction with python scripting and Selenium based web automation to streamline the generation of MD input files. The pipeline exploits CHARMM GUI's comprehensive web-based interface for preparing simulation-ready inputs for NAMD. By integrating Gemini's code generation and iterative refinement capabilities, simulation scripts are automatically written, executed, and revised to navigate CHARMM GUI, extract appropriate parameters, and produce the required NAMD input files. Post processing is performed using additional software to further refine the simulation outputs, thereby enabling a complete and largely hands free workflow. Our results demonstrate that this approach reduces setup time, minimizes manual errors, and offers a scalable solution for handling multiple protein systems in parallel. This automated framework paves the way for broader application of LLMs in computational structural biology, offering a robust and adaptable platform for future developments in simulation automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。