提出全局自适应蛋白分词法,提升结构生成与表示性能
Adaptive Protein Tokenization
- 采用逐级递增细节的全局分词思路,突破局部信息聚合限制
- 在重建、生成和分类任务中表现优于或持平现有方法
- 支持零样本蛋白缩减排程与亲和力优化,适合结构设计场景
分词是实现联合理解蛋白质序列、结构与功能的多模态模型的关键路径。现有蛋白质结构分词器通过局部邻域信息池化生成分词,限制了其在生成与表征任务中的表现。本文提出一种全局分词方法,使后续分词逐步增加全局表征的细节层次。该方法解决了基于局部分词的生成模型的多个问题:缓解误差累积、无需序列压缩操作即可获得嵌入表示,并支持任务定制的分词信息量调节。我们在重建、生成和表征任务上验证该方法,结果表明其性能达到或超过基于局部分词的现有模型。我们展示了自适应分词如何支持基于信息含量的推理准则,显著提升可设计性。在CATH分类任务上,对我们的分词序列进行非线性探测的表现优于其他分词器的对应表示。最后,我们证明该方法可支持零样本蛋白缩小和亲和力提升。
原文摘要 · Abstract (English)
Tokenization is a promising path to multi-modal models capable of jointly understanding protein sequences, structure, and function. Existing protein structure tokenizers create tokens by pooling information from local neighborhoods, an approach that limits their performance on generative and representation tasks. In this work, we present a method for global tokenization of protein structures in which successive tokens contribute increasing levels of detail to a global representation. This change resolves several issues with generative models based on local protein tokenization: it mitigates error accumulation, provides embeddings without sequence-reduction operations, and allows task-specific adaptation of a tokenized sequence's information content. We validate our method on reconstruction, generative, and representation tasks and demonstrate that it matches or outperforms existing models based on local protein structure tokenizers. We show how adaptive tokens enable inference criteria based on information content, which boosts designability. We validate representations generated from our tokenizer on CATH classification tasks and demonstrate that non-linear probing on our tokenized sequences outperforms equivalent probing on representations from other tokenizers. Finally, we demonstrate how our method supports zero-shot protein shrinking and affinity maturation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。