用专家知识增强AI,让化学合成路线评估更准更可解释。
Bridging Chemists and AI: An Expert-Augmented Framework for Interpretable Route Evaluation

- 结合机器学习与化学家经验,用树编辑距离训练模型
- 预测准确率60.2%,相关系数达0.78,远超旧方法
- 输出可理解的优/可接受/差三类评价,适合药物研发者
高效多步合成路线的选择是有机合成中的核心挑战,尤其在药物和工艺化学中,路线选择直接影响可行性、成本和开发效率。现有数据驱动评估系统常简化多目标设计问题,依赖专利路线等代理数据集,而非普适性标准。为此,我们提出一种专家增强的数据驱动评分框架,融合机器学习与化学家领域知识,实现量化与可解释的路线评估。基于DeepSets的模型利用参考路线与生成路线间的树编辑距离进行训练,并通过专家评价进行微调,输出定量分数和可解释的定性分类:良好、可行、差。该系统在类别预测上达到Spearman相关系数0.78和Pearson相关系数0.77,在分数预测上实现60.2%的顶级排名准确率,显著优于此前17.5%的基线水平。
原文摘要 · Abstract (English)
Selecting efficient multi-step synthetic routes is a central challenge in organic synthesis, particularly in medicinal and process chemistry, where route choice directly impacts feasibility, cost, and development efficiency. Data-driven assessment systems often oversimplify the multi-objective nature of synthesis design and rely on proxy datasets, such as patent routes, rather than universally grounded criteria. To address this, we introduce an expert-augmented, data-driven scoring framework that integrates machine learning with chemists' domain knowledge for both numerical and explainable route assessment. A DeepSets-based model is trained using tree edit distance between reference and machine-generated routes, and then fine-tuned with expert evaluations to produce both quantitative scores and interpretable qualitative categories: Good, Plausible, and Bad. The resulting system achieves a Spearman correlation coefficient of 0.78 and a Pearson correlation of 0.77 for category assessment prediction, and 60.2% top-1 ranking accuracy for score prediction, substantially outperforming the previous baseline of 17.5%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。