用可调替换矩阵提升蛋白属性预测,更高效且能融合结构信息。
Flexible Kernels for Protein Property Prediction

- 基于进化替换矩阵和局部线性设计灵活核函数,提升数据效率。
- 在结合结构信息后,多任务学习性能显著超越传统监督方法。
- 适合需要低样本量训练的蛋白设计与性质预测场景。
尽管在蛋白质设计中具有重要意义,但从稀疏实验数据中预测结合亲和力、热稳定性等蛋白属性仍面临巨大挑战。为此,我们提出一类序列核函数,利用进化替换矩阵及局部线性特性,证明其生成的高斯过程能高效建模蛋白属性景观,通常优于依赖基础模型嵌入的替代方法。进一步地,通过学习相当于结构感知的替换矩阵,我们的核函数可轻松整合来自基础模型的结构信息。实验表明,这些结构条件核函数在跨多个蛋白属性景观的多任务学习中表现优异,能显著超越局部监督学习方法。
原文摘要 · Abstract (English)
Despite its importance to applications in protein design, predicting protein properties like binding affinity and thermostability from sparse experimental data remains a significant challenge. Accordingly, we introduce a class of sequence kernels that exploit evolutionary substitution matrices as well as local linearity and demonstrate that the resulting Gaussian processes provide data-efficient models of protein property landscapes, frequently outperforming alternatives that rely on foundation model embeddings. Furthermore--by learning what are in effect structure-aware substitution matrices--we show that our kernels can readily incorporate structural information from foundation models. We demonstrate that these structure-conditioned kernels are well suited to multi-task learning across multiple protein property landscapes and can decisively outperform local supervised learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。