用电子流向建模反应机理,提升化学反应预测的可解释性。
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
- 将反应拆解为极性基元步骤,捕捉电子流动路径。
- 在PMechDB数据集上达到94.9%的前10准确率,路径恢复率达84.9%。
- 适合需要理解反应机理的药物研发与合成设计者。
准确预测化学反应对推动合成化学创新至关重要,广泛应用于医药、制造和农业领域。然而,反应预测复杂且耗时耗力。深度学习提供了高通量预测的可行方案,但现有模型多基于美国专利局数据集,将反应视为整体转化,缺乏可解释性与机理洞察。为此,我们提出PMechRP(极性机理反应预测器),基于PMechDB数据集训练模型,该数据集以极性基元步骤表示反应,体现电子流动与机理细节。为进一步扩展覆盖范围并提升泛化能力,我们通过组合生成扩充了PMechDB数据集。实验对比了多种架构:基于Transformer、图神经网络及两步Siamese模型。最佳方案为融合5个Chemformer模型与两步Siamese框架的混合模型,兼顾Transformer的精度,并利用两步网络过滤“非化学”产物。评估采用PMechDB测试集,另构建由有机化学教材提取的完整机理路径人类基准数据集。混合模型在PMechDB测试集上达94.9%的Top-10准确率,在路径数据集上目标恢复率达84.9%。
原文摘要 · Abstract (English)
Accurately predicting chemical reactions is essential for driving innovation in synthetic chemistry, with broad applications in medicine, manufacturing, and agriculture. At the same time, reaction prediction is a complex problem which can be both time-consuming and resource-intensive for chemists to solve. Deep learning methods offer an appealing solution by enabling high-throughput reaction prediction. However, many existing models are trained on the US Patent Office dataset and treat reactions as overall transformations: mapping reactants directly to products with limited interpretability or mechanistic insight. To address this, we introduce PMechRP (Polar Mechanistic Reaction Predictor), a system that trains machine learning models on the PMechDB dataset, which represents reactions as polar elementary steps that capture electron flow and mechanistic detail. To further expand model coverage and improve generalization, we augment PMechDB with a diverse set of combinatorially generated reactions. We train and compare a range of machine learning models, including transformer-based, graph-based, and two-step siamese architectures. Our best-performing approach was a hybrid model, which combines a 5-ensemble of Chemformer models with a two-step Siamese framework to leverage the accuracy of transformer architectures, while filtering away "alchemical" products using the two-step network predictions. For evaluation, we use a test split of the PMechDB dataset and additionally curate a human benchmark dataset consisting of complete mechanistic pathways extracted from an organic chemistry textbook. Our hybrid model achieves a top-10 accuracy of 94.9% on the PMechDB test set and a target recovery rate of 84.9% on the pathway dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。