通过优化残基特征表示,提升酶动力学参数预测精度。
KinForm: Kinetics Informed Feature Optimised Representation Models for Enzyme $k_{cat}$ and $K_{M}$ Prediction
- 基于结合位点概率加权池化,融合多模型残基嵌入。
- 在低序列相似性下性能提升显著,最高改善达15.3%。
- 适合需要强泛化能力的酶动力学研究者使用。
酶促反应的周转数($k_{cat}$)和米氏常数($K_{ ext{M}}$)对模拟酶活性至关重要,但实验数据规模与多样性有限。以往方法通常使用单个蛋白语言模型的均值池化残基嵌入表示蛋白质。本文提出KinForm,通过优化蛋白特征表示来提升动力学参数预测的准确性和泛化能力。KinForm融合来自多个残基级嵌入模型(Evolutionary Scale Modeling Cambrian、Evolutionary Scale Modeling 2、ProtT5-XL-UniRef50)的中间层特征,并基于每残基结合位点概率进行加权池化。为应对高维问题,采用主成分分析(PCA)对拼接特征降维,并通过基于相似性的过采样策略重平衡训练数据。在两个基准数据集上,KinForm优于基线方法,尤其在低序列相似性分组中表现更优。验证了结合位点概率池化、中间层选择、PCA与低相似性蛋白过采样的有效性。此外,去除折叠间序列重叠可提供更真实的泛化评估,建议作为标准测试策略。
原文摘要 · Abstract (English)
Kinetic parameters such as the turnover number ($k_{cat}$) and Michaelis constant ($K_{\mathrm{M}}$) are essential for modelling enzymatic activity but experimental data remains limited in scale and diversity. Previous methods for predicting enzyme kinetics typically use mean-pooled residue embeddings from a single protein language model to represent the protein. We present KinForm, a machine learning framework designed to improve predictive accuracy and generalisation for kinetic parameters by optimising protein feature representations. KinForm combines several residue-level embeddings (Evolutionary Scale Modeling Cambrian, Evolutionary Scale Modeling 2, and ProtT5-XL-UniRef50), taken from empirically selected intermediate transformer layers and applies weighted pooling based on per-residue binding-site probability. To counter the resulting high dimensionality, we apply dimensionality reduction using principal--component analysis (PCA) on concatenated protein features, and rebalance the training data via a similarity-based oversampling strategy. KinForm outperforms baseline methods on two benchmark datasets. Improvements are most pronounced in low sequence similarity bins. We observe improvements from binding-site probability pooling, intermediate-layer selection, PCA, and oversampling of low-identity proteins. We also find that removing sequence overlap between folds provides a more realistic evaluation of generalisation and should be the standard over random splitting when benchmarking kinetic prediction models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。