arXiv:2512.15133cs.CEcs.AI2025-12KDD被引 5

用连续结构符让蛋白语言模型同时理解序列和结构,效果媲美顶尖模型。

HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens

  • 在离散序列模型上叠加连续结构扩散头,统一处理两种模态。
  • 在无条件生成、逆折叠等任务中表现优秀,仅需不足十分之一算力。
  • 无需量化结构信息,适合追求高精度结构建模的研究者。

蛋白质具有固有的序列-结构双重性。尽管蛋白质序列数据丰富且易于表示为离散标记,推动了蛋白语言模型(pLMs)的发展,但如何有效整合连续结构知识仍是关键挑战。现有方法通常将结构离散化以适配语言建模框架,导致精细信息丢失,限制多模态模型性能。本文提出一种混合扩散蛋白语言模型HD-Prot,通过在离散pLM之上嵌入连续值扩散头,实现对离散与连续标记的无缝联合建模。它利用统一吸收扩散过程捕捉跨模态的词间依赖,并通过类别预测建模序列分布,连续扩散建模结构分布。大量实验表明,HD-Prot在无条件序列-结构联合生成、基序支架构建、蛋白质结构预测及逆折叠任务中表现优异。尽管计算资源有限(扩展微调预算不足顶级模型的十分之一),其性能仍可比肩先进多模态pLMs。该方法证明了在统一架构中同时估计分类与连续分布的可行性,为多模态蛋白语言模型提供了新方向。

原文摘要 · Abstract (English)

Proteins inherently possess a consistent sequence-structure duality. The abundance of protein sequence data, which can be readily represented as discrete tokens, has driven fruitful developments in protein language models (pLMs). A key remaining challenge, however, is how to effectively integrate continuous structural knowledge into pLMs. Current methods often discretize protein structures to accommodate the language modeling framework, which inevitably results in the loss of fine-grained information and limits the performance potential of multimodal pLMs. In this paper, we argue that such concerns can be circumvented: a sequence-based pLM can be extended to incorporate the structure modality through continuous tokens, i.e., high-fidelity protein structure latents that avoid vector quantization. Specifically, we propose a hybrid diffusion protein language model, HD-Prot, which embeds a continuous-valued diffusion head atop a discrete pLM, enabling seamless operation with both discrete and continuous tokens for joint sequence-structure modeling. It captures inter-token dependencies across modalities through a unified absorbing diffusion process, and estimates per-token distributions via categorical prediction for sequences and continuous diffusion for structures. Extensive results demonstrate that HD-Prot achieves competitive performance in unconditional sequence-structure co-generation, motif-scaffolding, protein structure prediction, and inverse folding tasks. Furthermore, our method can perform on par with state-of-the-art multimodal pLMs, despite being developed under limited computational resources (i.e., less than one-tenth the budget for modality extension fine-tuning). It highlights the viability of simultaneously estimating categorical and continuous distributions within a unified language model architecture, offering a promising alternative direction for multimodal pLMs.

蛋白语言模型结构建模扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。