用扩散模型提升蛋白质序列表征,增强判别能力
Discriminative protein sequence modelling with Latent Space Diffusion
- 将蛋白质序列编码器与潜空间扩散模型结合,学习多层级表征
- 两种扩散模型均比掩码语言模型基线具备更强判别力
- 适合需要高区分度蛋白表征的下游预测任务
我们提出一种蛋白质序列表征学习框架,将任务分解为流形学习与分布建模。具体地,设计了一种潜空间扩散架构,结合蛋白质序列自编码器与在潜空间运行的去噪扩散模型。由此获得一个单参数族的表征,以及自编码器本身的潜空间表示。我们提出了两种自编码器结构:一种强制相同氨基酸类型在潜空间中同分布(同质模型),另一种采用基于噪声的掩码变体(异质模型)。以掩码语言建模学习的潜空间为基线,在多种蛋白质属性预测任务上评估判别能力。结果表明:两种扩散模型训练出的表征判别力均优于掩码语言模型基线,但未达到掩码语言模型嵌入自身的性能。
原文摘要 · Abstract (English)
We explore a framework for protein sequence representation learning that decomposes the task between manifold learning and distributional modelling. Specifically we present a Latent Space Diffusion architecture which combines a protein sequence autoencoder with a denoising diffusion model operating on its latent space. We obtain a one-parameter family of learned representations from the diffusion model, along with the autoencoder's latent representation. We propose and evaluate two autoencoder architectures: a homogeneous model forcing amino acids of the same type to be identically distributed in the latent space, and an inhomogeneous model employing a noise-based variant of masking. As a baseline we take a latent space learned by masked language modelling, and evaluate discriminative capability on a range of protein property prediction tasks. Our finding is twofold: the diffusion models trained on both our proposed variants display higher discriminative power than the one trained on the masked language model baseline, none of the diffusion representations achieve the performance of the masked language model embeddings themselves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。