让生物学家用自然语言操控深度学习,自动设计蛋白质。
AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering
- 用大模型+自动化机器学习,让非编程人员也能操作蛋白设计
- 在真实任务中比传统方法性能显著提升
- 适合无计算背景的生物科研人员快速上手
蛋白质工程对生物医学应用至关重要,但传统方法效率低且资源消耗大。尽管深度学习(DL)展现出潜力,但其训练与应用仍对缺乏计算专长的生物学家构成挑战。为此,我们提出AutoProteinEngine(AutoPE),一个基于大语言模型(LLMs)的多模态自动化机器学习(AutoML)代理框架。AutoPE创新性地使无深度学习背景的生物学家可通过自然语言与深度学习模型交互,降低蛋白工程门槛。该框架独特融合大模型与AutoML,实现蛋白序列与图结构模态的模型选择、超参数自动优化及蛋白数据库的自动数据获取。我们在两个真实蛋白工程任务中评估了AutoPE,结果表明其性能显著优于传统零样本和手动微调方法。通过弥合深度学习与生物学家领域知识之间的鸿沟,AutoPE使研究人员无需大量编程即可利用深度学习。代码已开源:https://github.com/tsynbio/AutoPE。
原文摘要 · Abstract (English)
Protein engineering is important for biomedical applications, but conventional approaches are often inefficient and resource-intensive. While deep learning (DL) models have shown promise, their training or implementation into protein engineering remains challenging for biologists without specialized computational expertise. To address this gap, we propose AutoProteinEngine (AutoPE), an agent framework that leverages large language models (LLMs) for multimodal automated machine learning (AutoML) for protein engineering. AutoPE innovatively allows biologists without DL backgrounds to interact with DL models using natural language, lowering the entry barrier for protein engineering tasks. Our AutoPE uniquely integrates LLMs with AutoML to handle model selection for both protein sequence and graph modalities, automatic hyperparameter optimization, and automated data retrieval from protein databases. We evaluated AutoPE through two real-world protein engineering tasks, demonstrating substantial performance improvements compared to traditional zero-shot and manual fine-tuning approaches. By bridging the gap between DL and biologists' domain expertise, AutoPE empowers researchers to leverage DL without extensive programming knowledge. Our code is available at https://github.com/tsynbio/AutoPE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。