将蛋白质序列作为原生模态融入大模型,实现理解与设计一体化。
AMix-2: Establishing Protein as a Native Modality in Large Language Models

- 用共享词元空间统一编码文本和蛋白序列,支持联合推理与生成。
- 采用分块扩散建模,可灵活调整生成顺序,性能优于传统自回归模型。
- 提出新基准ProteinArena,适合评估真实场景下的通用性与泛化能力。
我们提出AMix-2,一种将蛋白质作为原生模态的大语言模型基础模型,统一了蛋白质理解与序列设计任务。其核心思想包括:(1)构建统一的蛋白-文本表征框架,将自然语言与蛋白序列嵌入共享词元空间,使单一模型能同时完成生物推理与条件生成,避免依赖下游任务专用模型;(2)采用分块扩散语言建模范式,结合块间因果生成、块内双向上下文与迭代优化机制,更贴合蛋白质内在结构特性,优于严格左到右的因子分解方式。为在真实泛化设置下评估蛋白基础模型,我们引入ProteinArena——一个涵盖时间感知与同源感知协议的综合性基准,覆盖多种理解与设计任务,并包含经典生物信息工具、蛋白专用模型及大语言模型等基线。在ProteinArena上,AMix-2优于前沿大语言模型,并达到与任务特化蛋白模型相当的性能。控制实验表明,基于扩散的范式普遍优于自回归模型,凸显灵活生成顺序对蛋白序列生成的优势。我们已开源AMix-2与ProteinArena,推动蛋白基础模型的开放研究。
原文摘要 · Abstract (English)
We present AMix-2, a protein-text foundation model that establishes protein as a native modality in large language models (LLMs), unifying protein understanding and sequence design within a single foundation model. AMix-2 is built upon two key ideas: (1) a unified protein-text formulation that embeds natural language and protein sequence in a shared token space, enabling one model to perform biological reasoning and conditional design instead of separate downstream task-specialized models; and (2) a block-wise diffusion language modeling backbone that combines causal generation across blocks with bidirectional context and iterative refinement within blocks. This scheme better matches the intrinsic nature of proteins than a strict left-to-right factorization. To evaluate protein foundation models under realistic generalization settings, we further introduce ProteinArena, a comprehensive benchmark with time-aware and homology-aware protocols across various understanding and design tasks, and with baselines covering classical bioinformatics tools, protein-specialized models and LLMs. On ProteinArena, AMix-2 outperforms frontier LLMs and demonstrates competitive performance to task-specific protein models. Controlled experiments further show that the diffusion-based paradigm generally surpasses its autoregressive counterpart, highlighting the advantage of flexible generation order for protein sequences. We release both AMix-2 and ProteinArena to facilitate open research in protein foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。