融合序列、底物与环境信息预测酶促反应速率,提升准确性与可解释性。
Multimodal Regression for Enzyme Turnover Rates Prediction
- 整合蛋白序列、底物结构和环境因子,用多模态模型联合建模
- 在多个数据集上超越现有方法,实现更精准的酶促速率预测
- 通过柯尔莫哥洛夫网络显式学习数学公式,增强结果可解释性
酶促反应速率是酶动力学中的关键参数,反映酶的催化效率。然而,由于实验测量成本高、难度大,大多数生物体中的酶促反应速率数据仍然稀缺。为填补这一空白,我们提出一种多模态框架,通过整合酶序列、底物结构和环境因素来预测酶促反应速率。模型结合预训练语言模型与卷积神经网络提取蛋白质序列特征,利用图神经网络捕捉底物分子的信息表示,并引入注意力机制增强酶与底物表征间的交互。此外,我们采用柯尔莫哥洛夫-阿诺德网络进行符号回归,显式学习描述酶促反应速率的数学表达式,从而实现可解释且高精度的预测。大量实验表明,该框架在多个基准上优于传统及前沿深度学习方法。本研究为酶动力学研究提供有力工具,对酶工程、生物技术和工业生物催化具有应用前景。
原文摘要 · Abstract (English)
The enzyme turnover rate is a fundamental parameter in enzyme kinetics, reflecting the catalytic efficiency of enzymes. However, enzyme turnover rates remain scarce across most organisms due to the high cost and complexity of experimental measurements. To address this gap, we propose a multimodal framework for predicting the enzyme turnover rate by integrating enzyme sequences, substrate structures, and environmental factors. Our model combines a pre-trained language model and a convolutional neural network to extract features from protein sequences, while a graph neural network captures informative representations from substrate molecules. An attention mechanism is incorporated to enhance interactions between enzyme and substrate representations. Furthermore, we leverage symbolic regression via Kolmogorov-Arnold Networks to explicitly learn mathematical formulas that govern the enzyme turnover rate, enabling interpretable and accurate predictions. Extensive experiments demonstrate that our framework outperforms both traditional and state-of-the-art deep learning approaches. This work provides a robust tool for studying enzyme kinetics and holds promise for applications in enzyme engineering, biotechnology, and industrial biocatalysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。