首个可处理蛋白质构象动态的分词器,让语言模型理解蛋白运动
ENSEMBITS: an alphabet of protein conformational ensembles

- 基于残差VQ-VAE与帧蒸馏目标,从多构象数据中提取动态特征
- 在残差运动幅度预测上超越所有现有方法,零样本突变影响预测表现优异
- 仅用少量预训练数据即可媲美静态分词器,适合动态结构建模与设计
蛋白质结构分词器(PSTs)是蛋白质语言建模、功能预测和进化分析的核心工具。然而,现有PSTs仅捕捉静态结构的局部几何信息,忽略了由蛋白质集合揭示的相关运动和构象状态。本文提出Ensembits,首个用于蛋白质构象集合的分词器。它解决了动态数据分词的三大挑战:跨构象的有意义几何描述、可变大小集合的置换不变编码,以及动态数据稀疏性问题。通过在大规模分子动力学语料上使用残差VQ-VAE和帧蒸馏目标进行训练,Ensembits在均方根波动(RMSF)预测上优于所有相关方法,并在残基级运动幅度的条件化ANOVA测试中成为最强独立结构分词器。尽管预训练数据远少于静态分词器,其在酶促反应(EC)、基因本体(GO)、结合位点/亲和力预测及零样本突变效应预测任务上仍达到或超过后者表现。值得注意的是,蒸馏目标使Ensembits能仅凭单一预测结构生成动态分词,有效缓解动态数据稀疏问题。随着领域从静态结构预测转向集合生成,Ensembits为将动态信息引入蛋白质语言建模与设计提供了离散词汇基础。
原文摘要 · Abstract (English)
Protein structure tokenizers (PSTs) are workhorses in protein language modeling, function prediction, and evolutionary analysis. However, existing PSTs only capture local geometry of static structures, and miss the correlated motions and alternative conformational states revealed by protein ensembles. Here we introduce Ensembits, the first tokenizer of protein conformational ensembles. Ensembits address challenges inherent to tokenizing dynamics: deriving informative geometric descriptors across conformations, permutation-invariance encoding of variable-size ensembles, and conquering sparsity in dynamics data. Trained with a Residual VQ-VAE using a frame distillation objective on a large molecular dynamics corpus, Ensembits outperforms all related methods on RMSF prediction, and is the strongest standalone structural tokenizer on an token-conditioned ANOVA test on per-residue motion amplitude. Ensembits further matches or exceeds static tokenizers on EC, GO, binding site/affinity prediction, and zero-shot mutation-effect prediction despite using far less pretraining data. Notably, the distillation objective enables Ensembits to predict dynamics token from one single predicted structure, which alleviates dynamics data sparsity. As the field moves from static structure prediction toward ensemble generation, Ensembits offer the discrete vocabulary needed to bring dynamics into protein language modeling and design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。