arXiv:2505.05616cs.AIcs.LG2025-05被引 1

用大模型预测酶反应,提升生物催化研发效率

Leveraging Large Language Models for enzymatic reaction prediction and characterization

  • 用Llama-3.1大模型做酶反应预测,支持正向与逆向合成
  • 多任务学习使反应预测准确率显著提升,尤其在数据少时表现好
  • 适合做酶工程、药物发现的研究者参考

酶反应预测对生物催化、代谢工程和药物发现至关重要,但仍是复杂且耗资源的任务。大型语言模型(LLMs)在科学领域展现出强大潜力,具备知识泛化、复杂结构推理和上下文学习能力。本研究系统评估了Llama-3.1系列(8B和70B)在三大核心生化任务中的表现:酶委员会编号预测、正向合成与逆向合成。比较了单任务与多任务学习策略,采用参数高效微调(LoRA适配器)。还考察了不同数据规模下的表现,探索其在低数据场景的适应性。结果表明,微调后的LLM能有效捕捉生化知识,多任务学习通过共享酶学信息提升了正向与逆向合成预测性能。同时发现层级式EC分类存在挑战,指明了未来改进方向。

原文摘要 · Abstract (English)

Predicting enzymatic reactions is crucial for applications in biocatalysis, metabolic engineering, and drug discovery, yet it remains a complex and resource-intensive task. Large Language Models (LLMs) have recently demonstrated remarkable success in various scientific domains, e.g., through their ability to generalize knowledge, reason over complex structures, and leverage in-context learning strategies. In this study, we systematically evaluate the capability of LLMs, particularly the Llama-3.1 family (8B and 70B), across three core biochemical tasks: Enzyme Commission number prediction, forward synthesis, and retrosynthesis. We compare single-task and multitask learning strategies, employing parameter-efficient fine-tuning via LoRA adapters. Additionally, we assess performance across different data regimes to explore their adaptability in low-data settings. Our results demonstrate that fine-tuned LLMs capture biochemical knowledge, with multitask learning enhancing forward- and retrosynthesis predictions by leveraging shared enzymatic information. We also identify key limitations, for example challenges in hierarchical EC classification schemes, highlighting areas for further improvement in LLM-driven biochemical modeling.

酶反应预测大模型生物催化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。