arXiv:2410.07919cs.CLq-bio.BM2024-10被引 18

用自然语言指导设计蛋白质和药物分子,提升精准度与效率。

Advancing biomolecular understanding and design following human instructions

  • 通过多模态对齐,让大模型理解人类语言指令
  • 生成药物结合亲和力提升10%,酶底物匹配率达70.4%
  • 适合生物制药与酶工程领域的研究人员使用

理解与设计生物分子(如蛋白质和小分子)是推动药物发现、合成生物学和酶工程的关键。近年来,人工智能在生物分子预测与设计方面取得突破性进展,但其计算能力与研究者直观目标之间仍存在显著差距,尤其是在利用自然语言衔接复杂任务与人类意图方面。大语言模型虽有潜力解析人类意图,但在生物分子研究中的应用仍受限于专业知识要求高、多模态数据融合难、自然语言与分子语义对齐不足等问题。为此,我们提出 InstructBioMol,一个通过自然语言、分子与蛋白质之间的全向对齐,实现任意形式输入输出的大型语言模型。该模型可整合多模态生物分子输入,支持研究者以自然语言描述设计目标,并生成符合精确生物学需求的生物分子输出。实验表明,InstructBioMol 能准确理解并执行人类指令,生成的药物分子结合亲和力提升10%,设计的酶-底物配对预测得分达70.4。这展示了其在真实生物分子研究中的转化潜力。代码已开源:https://github.com/HICAI-ZJU/InstructBioMol。

原文摘要 · Abstract (English)

Understanding and designing biomolecules, such as proteins and small molecules, is central to advancing drug discovery, synthetic biology and enzyme engineering. Recent breakthroughs in artificial intelligence have revolutionized biomolecular research, achieving remarkable accuracy in biomolecular prediction and design. However, a critical gap remains between artificial intelligence's computational capabilities and researchers' intuitive goals, particularly in using natural language to bridge complex tasks with human intentions. Large language models have shown potential to interpret human intentions, yet their application to biomolecular research remains nascent due to challenges including specialized knowledge requirements, multimodal data integration, and semantic alignment between natural language and biomolecules. To address these limitations, we present InstructBioMol, a large language model designed to bridge natural language and biomolecules through a comprehensive any-to-any alignment of natural language, molecules and proteins. This model can integrate multimodal biomolecules as the input, and enable researchers to articulate design goals in natural language, providing biomolecular outputs that meet precise biological needs. Experimental results demonstrate that InstructBioMol can understand and design biomolecules following human instructions. In particular, it can generate drug molecules with a 10% improvement in binding affinity and design enzymes that achieve an enzyme-substrate pair prediction score of 70.4. This highlights its potential to transform real-world biomolecular research. The code is available at https://github.com/HICAI-ZJU/InstructBioMol.

生物分子设计大模型自然语言药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。