通过多层级对齐提升酶与反应匹配的精准度。
Multi-Alignment Contrastive Learning for Enzyme--Reaction Retrieval
- 引入跨域与域内双重对齐的对比学习框架。
- 在EnzymeMap上早期识别率显著优于基线模型。
- 适合酶功能预测与生物催化设计场景。
识别催化特定生化反应的酶是计算酶发现与生物催化剂设计的关键步骤。现有表示学习方法将此问题建模为酶-反应匹配任务,将配对的酶与反应嵌入共享空间。然而,大多数方法仅依赖成对酶-反应监督,未充分利用反应集或酶家族内的关系。本文提出一种多对齐对比学习框架,联合建模酶与反应间的跨域兼容性以及由功能注释引出的域内关系。此外,基于Gromov-Wasserstein的正则化目标促使学习到的酶与反应表示空间保持几何一致性。通过结合成对催化监督与高阶关系对齐,模型同时捕捉直接酶-反应关联与更广泛的的功能组织结构。我们在酶虚拟筛选与双向酶-反应检索任务上进行评估。在EnzymeMap数据集上,相较于强对比基线,在BEDROC与富集因子指标下实现更好的早期识别性能。在ReactZyme数据集上,方法在基于时间、酶相似性和反应相似性的分割中均取得一致提升,表明其对未见酶和未见反应具有鲁棒性。消融实验进一步显示,域内对齐、功能监督与几何正则项均对性能提升有贡献。结果表明,建模多种对齐形式可有效提升对比检索模型在酶发现、反应注释及相关计算生物学应用中的表现。
原文摘要 · Abstract (English)
Identifying enzymes that catalyze target biochemical reactions is a key step in computational enzyme discovery and biocatalyst design. Recent representation-learning methods formulate this problem as enzyme--reaction matching, where paired enzymes and reactions are embedded into a shared space. However, most existing approaches primarily rely on pairwise enzyme--reaction supervision and make limited use of the relationships within reaction sets or enzyme families. This work introduces a multi-alignment contrastive learning framework for biochemical retrieval. The framework jointly models cross-domain compatibility between enzymes and reactions and within-domain relationships induced by functional annotations. In addition, a Gromov--Wasserstein-inspired regularization objective encourages geometric consistency between the learned enzyme and reaction representation spaces. By combining pairwise catalytic supervision with higher-order relational alignment, the model captures both direct enzyme--reaction associations and broader functional organization. We evaluate the approach on enzyme virtual screening and bidirectional enzyme--reaction retrieval tasks. Experiments on EnzymeMap show improved early-recognition performance under BEDROC and enrichment-factor metrics compared with strong contrastive baselines. On ReactZyme, the method achieves consistent gains across time-based, enzyme-similarity, and reaction-similarity splits, demonstrating robustness to unseen enzymes and unseen reactions. Ablation studies further indicate that within-domain alignment, functional supervision, and the geometric regularization term each contribute to the observed improvements. These results suggest that modeling multiple forms of alignment can improve contrastive retrieval models for enzyme discovery, reaction annotation, and related computational biology applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。