开源蛋白语言模型助力功能预测与酶设计,降低计算门槛。
Open-Source Protein Language Models for Function Prediction and Protein Design
- 将蛋白语言模型集成至DeepChem框架,提升可访问性。
- 在多个任务上表现合理,生成降解塑料的酶候选物。
- 适合资源有限的研究者开展合成生物学与环保研究。
蛋白语言模型(PLMs)在理解蛋白序列方面展现出潜力,推动了功能预测和蛋白质工程的发展。然而,从头训练这些模型需要大量计算资源,限制了其普及。为此,我们将PLM集成到DeepChem这一开源计算生物化学框架中,提供更易获取的蛋白相关任务平台。在多个蛋白预测任务上评估模型性能,结果表明其在基准测试中表现合理。此外,我们探索利用模型嵌入和潜在空间操控技术生成降解塑料的酶候选物。尽管结果仍需进一步优化,但该方法为未来酶设计研究奠定了基础。本研究旨在使PLM在合成生物学和环境可持续性等领域更易被使用,尤其适用于计算资源有限的研究者。
原文摘要 · Abstract (English)
Protein language models (PLMs) have shown promise in improving the understanding of protein sequences, contributing to advances in areas such as function prediction and protein engineering. However, training these models from scratch requires significant computational resources, limiting their accessibility. To address this, we integrate a PLM into DeepChem, an open-source framework for computational biology and chemistry, to provide a more accessible platform for protein-related tasks. We evaluate the performance of the integrated model on various protein prediction tasks, showing that it achieves reasonable results across benchmarks. Additionally, we present an exploration of generating plastic-degrading enzyme candidates using the model's embeddings and latent space manipulation techniques. While the results suggest that further refinement is needed, this approach provides a foundation for future work in enzyme design. This study aims to facilitate the use of PLMs in research fields like synthetic biology and environmental sustainability, even for those with limited computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。