无需训练,用随机注意力从少量蛋白序列生成新序列。
Training-Free Generation of Protein Sequences from Small Family Alignments via Stochastic Attention
- 基于注意力残差构建无训练采样器,直接从对齐序列中生成
- 生成序列在组成、新颖性和结构合理性上均优于现有方法
- 仅需主成分分析即可自动设定参数,适合小家族蛋白
生成符合蛋白家族统计特性的新序列通常需要在数千至数百万条序列上训练深度生成模型。然而大多数蛋白家族规模很小:中位Pfam种子比对仅含22条序列,此时学习模型易过拟合或崩溃。本文提出随机注意力(SA),一种无需训练的采样器,将现代霍普菲尔德能量视为玻尔兹曼分布,并通过朗之万动力学采样。其得分函数为单个softmax注意力操作的残差,无需训练得分网络、预训练数据或GPU。在涵盖37至420条序列、23至262个残基的八个Pfam家族中,SA生成的序列具有低组成偏差、高新颖性及结构合理性(由ESMFold和AlphaFold2验证)。相比谱型隐马尔可夫模型(HMMs)、EvoDiff和多序列比对变换器(MSA Transformer),SA是唯一同时实现低组成偏差、真实新颖性且序列相似度保持在家族最近邻范围内的方法;其余方法则偏离该范围或生成近似拷贝。关键逆温度可仅由主成分分析维度预测,实现从种子比对到完全自动化的操作。在两个具备深突变扫描数据的领域中,SA生成的替换在实验耐受突变中显著富集,超出位置匹配的零模型,且独立语言模型(ESM2-650M)对其评分处于自然范围内。随机注意力因此使训练无关的序列生成成为可能,覆盖深度学习难以触及的小型蛋白家族。
原文摘要 · Abstract (English)
Generating novel protein sequences that respect a family's statistical constraints typically requires training deep generative models on thousands to millions of examples. Yet most protein families are small: the median Pfam seed alignment contains only 22 sequences, a regime where learned models overfit or collapse. We propose \emph{stochastic attention} (SA), a training-free sampler that treats the modern Hopfield energy over stored sequences as a Boltzmann distribution and draws samples via Langevin dynamics. The score function is the residual of a single softmax attention operation, eliminating the need for a trained score network, pretraining data, or graphics processing units (GPUs). Across eight Pfam families spanning 37 to 420 sequences and 23 to 262 residues, SA generates sequences with low composition divergence, novelty, and structural plausibility supported by ESMFold and AlphaFold2. Compared with profile hidden Markov models (HMMs), EvoDiff, and the multiple sequence alignment (MSA) Transformer, SA is the only tested method to simultaneously achieve low composition divergence, genuine novelty, and sequence identity within each family's nearest-neighbor identity range; the others drift outside this range or produce near-copies. The critical inverse temperature is predicted from principal component analysis (PCA) dimensionality alone, enabling fully automatic operation from a seed alignment. In two domains with deep mutational scanning data, SA-generated substitutions are enriched for experimentally tolerated mutations beyond a position-matched null, and an independent language model (ESM2-650M) scores them within the natural range. Stochastic attention thus opens training-free sequence generation to the long tail of protein families too small for deep learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。