用多模态模型分阶段预测酶催化效率,提升对底物识别和构象适应的建模能力。
Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation
- 分两阶段建模:先融合底物信息捕获特异性,再根据活性位点结构路由到专用专家
- 在三个动力学指标上均优于传统方法,尤其在分布外数据表现更稳健
- 适合需要精准酶动力学预测的研究者,如药物设计与蛋白质工程
预测酶动力学参数可量化酶在特定生化条件下催化特定底物的效率。典型参数如周转数(k_cat)、米氏常数(K_m)和抑制常数(K_i),共同依赖于酶序列、底物化学性质以及结合过程中活性位点的构象适应。现有学习流程常将此简化为酶与底物的静态兼容性问题,通过浅层融合表示并回归单一数值,忽略了催化过程的阶段性特征,即底物识别与构象适应。为此,本文将动力学预测重构为分阶段的多模态条件建模问题,提出酶-反应桥接适配器(ERBA),通过微调将跨模态信息注入蛋白质语言模型(PLMs),同时保留其生物化学先验。ERBA采用双阶段条件化:分子识别交叉注意力(MRCA)首先将底物信息注入酶表征以捕捉特异性;几何感知专家混合(G-MoE)则整合活性位点结构,将样本路由至口袋专用专家以反映诱导契合。为保持语义一致性,酶-底物分布对齐(ESDA)在再生核希尔伯特空间中强制约束PLM流形内的分布一致性。在三个动力学终点及多种PLM主干网络上的实验表明,ERBA持续优于仅使用序列或浅层融合的基线方法,且在分布外数据上表现更强,为可扩展的动力学预测提供了生物学合理路径,并可拓展加入辅因子、突变和时序结构线索。
原文摘要 · Abstract (English)
Predicting enzyme kinetic parameters quantifies how efficiently an enzyme catalyzes a specific substrate under defined biochemical conditions. Canonical parameters such as the turnover number ($k_\text{cat}$), Michaelis constant ($K_\text{m}$), and inhibition constant ($K_\text{i}$) depend jointly on the enzyme sequence, the substrate chemistry, and the conformational adaptation of the active site during binding. Many learning pipelines simplify this process to a static compatibility problem between the enzyme and substrate, fusing their representations through shallow operations and regressing a single value. Such formulations overlook the staged nature of catalysis, which involves both substrate recognition and conformational adaptation. In this regard, we reformulate kinetic prediction as a staged multimodal conditional modeling problem and introduce the Enzyme-Reaction Bridging Adapter (ERBA), which injects cross-modal information via fine-tuning into Protein Language Models (PLMs) while preserving their biochemical priors. ERBA performs conditioning in two stages: Molecular Recognition Cross-Attention (MRCA) first injects substrate information into the enzyme representation to capture specificity; Geometry-aware Mixture-of-Experts (G-MoE) then integrates active-site structure and routes samples to pocket-specialized experts to reflect induced fit. To maintain semantic fidelity, Enzyme-Substrate Distribution Alignment (ESDA) enforces distributional consistency within the PLM manifold in a reproducing kernel Hilbert space. Experiments across three kinetic endpoints and multiple PLM backbones, ERBA delivers consistent gains and stronger out-of-distribution performance compared with sequence-only and shallow-fusion baselines, offering a biologically grounded route to scalable kinetic prediction and a foundation for adding cofactors, mutations, and time-resolved structural cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。