用混合变分自编码器设计能模仿蛋白质功能的随机共聚物。
DeepRHP: A Hybrid Variational Autoencoder for Designing Random Heteropolymers as Protein Mimics

- 结合经典VAE与特征增强VAE,让潜在空间同时捕捉化学特征和序列模式。
- 预测出可稳定膜蛋白(如水通道蛋白Z)的单体组合,结果与已有文献吻合。
- 适合生物材料设计、蛋白质模拟方向的研究者参考。
合成随机共聚物(RHPs)由预定义的单体组成,是设计类蛋白材料的有效途径。若设计得当,RHPs可模仿蛋白质的行为与功能。为此,我们开发了DeepRHP,一种基于半监督框架的改进型变分自编码器(VAE)模型。通过在传统VAE基础上增加一个基于特征的额外VAE,DeepRHP强制潜在空间同时表征关键化学特征与单个RHP序列模式,实现灵活融合任意相关特征。我们验证了该方法的有效性:成功预测出可使膜蛋白(如Aquaporin Z)在非天然环境中稳定的单体组成,并通过已有实验数据交叉验证。模型预测与真实RHP功能高度一致,表明混合自编码架构在指导RHP设计用于蛋白质及其他生物分子方面具有巨大潜力。
原文摘要 · Abstract (English)
Synthetic random heteropolymers (RHPs), consisting of a predefined set of monomers, offer an approach toward the design of protein-like materials. These RHPs, if designed appropriately, can mimic protein behavior and function. As such, there is a need for computational tools to efficiently guide RHP design. We bridge this gap by developing DeepRHP, a modified variational autoencoder (VAE) model under a semi-supervised framework. By equipping a classical VAE with an additional feature-based VAE, DeepRHP forces the latent space to capture structures of critical chemical features as well as individual RHP sequence patterns. In this sense, our method is versatile by allowing any relevant features to be incorporated in a hybrid manner. We demonstrate the effectiveness of DeepRHP by suggesting potential monomer compositions that stabilize membrane proteins (e.g. Aquaporin Z) in non-native environments and cross-validating our prediction with published results. The concordance between our model and true RHP function suggests strong potential in utilizing hybrid autoencoder architectures to guide RHP design for proteins and other biological compounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。