arXiv:2503.08179q-bio.BMcs.AI2025-03被引 11

用统一编码让大模型理解蛋白质结构,实现功能预测与设计

ProtTeX: Structure-In-Context Reasoning and Editing of Proteins with Large Language Models

  • 将序列、结构、文本统一编码到离散空间,支持多模态输入输出
  • 功能预测准确率提升一倍,超越专业模型,生成高质量构象
  • 无需特殊训练即可让通用大模型处理蛋白质任务,适合生物研发人员

大语言模型在小分子研究中取得显著进展,主要得益于有效的分子分词策略。然而,在蛋白质科学中,仅使用氨基酸序列作为分词方式,缺乏结构感知能力,限制了模型对蛋白质的全面理解与多模态生成。为此,我们提出 ProtTeX 框架,将蛋白质序列、结构和文本信息统一映射到离散空间,通过标准的 Next-Token Prediction 训练范式实现联合训练。该方法使通用大模型可通过文本输入感知并处理蛋白质结构,利用结构作为推理中间变量,并通过文本输出生成或编辑结构。实验表明,模型在蛋白质功能预测上表现优异,准确率较现有最优领域模型提升两倍;同时可生成高质量构象并实现定制化蛋白设计。首次证明,采用标准大模型训练与推理流程,解码器仅有的大模型可有效应对多样化的蛋白质任务。

原文摘要 · Abstract (English)

Large language models have made remarkable progress in the field of molecular science, particularly in understanding and generating functional small molecules. This success is largely attributed to the effectiveness of molecular tokenization strategies. In protein science, the amino acid sequence serves as the sole tokenizer for LLMs. However, many fundamental challenges in protein science are inherently structure-dependent. The absence of structure-aware tokens significantly limits the capabilities of LLMs for comprehensive biomolecular comprehension and multimodal generation. To address these challenges, we introduce a novel framework, ProtTeX, which tokenizes the protein sequences, structures, and textual information into a unified discrete space. This innovative approach enables joint training of the LLM exclusively through the Next-Token Prediction paradigm, facilitating multimodal protein reasoning and generation. ProtTeX enables general LLMs to perceive and process protein structures through sequential text input, leverage structural information as intermediate reasoning components, and generate or manipulate structures via sequential text output. Experiments demonstrate that our model achieves significant improvements in protein function prediction, outperforming the state-of-the-art domain expert model with a twofold increase in accuracy. Our framework enables high-quality conformational generation and customizable protein design. For the first time, we demonstrate that by adopting the standard training and inference pipelines from the LLM domain, ProtTeX empowers decoder-only LLMs to effectively address diverse spectrum of protein-related tasks.

蛋白质生成大模型结构感知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。