解决化学反应数据缺失问题,提升反应补全的准确性与可靠性
CompleteRXN: Toward Completing Open Chemical Reaction Databases

- 构建真实缺失场景下的大规模反应补全基准数据集
- 新模型CRB在随机数据上达99.20%准确率,极端分布外仍保持91.12%
- 揭示现有方法在真实数据上的性能下降,强调实际应用挑战
USPTO等化学反应数据集普遍存在副产物、共反应物及计量系数缺失的问题,严重影响下游应用的可用性与可靠性。本文提出CompleteRXN,一个面向真实缺失条件的大规模监督基准,通过将USPTO记录映射到精心校准的机理反应,构建了对齐的不完整与原子平衡反应数据集。评估了多种基线模型,包括一种带约束解码的新编码器-解码器模型CRB和近期算法方法SynRBL。在CompleteRXN基准上,CRB在随机划分下达到99.20%等价准确率,在极端分布外划分下仍达91.12%。SynRBL生成大量平衡且化学合理的补全结果,但基准测试准确率较低。所有方法性能随缺失程度增加而下降,尤其在未校准的USPTO全集上表现显著退化,凸显基准性能与实际鲁棒性之间的差距,为未来研究提供方向。
原文摘要 · Abstract (English)
Chemical reaction datasets such as USPTO suffer from substantial incompleteness, frequently missing byproducts, co-reactants, and stoichiometric coefficients. This limits their applicability and reliability in downstream applications. Here, we introduce CompleteRXN, a large-scale supervised benchmark for reaction completion under realistic missing-data conditions. We construct a dataset of aligned incomplete and atom-balanced reactions by mapping USPTO records to curated mechanistic reactions. We evaluate representative baselines, including a novel encoder-decoder reaction completion model with constrained decoding, the Constrained Reaction Balancer (CRB), and a recent algorithmic method, SynRBL. On our CompleteRXN benchmark, the CRB achieves high performance across splits of increasing difficulty, reaching 99.20% equivalence accuracy on the random split and 91.12% on the extreme out-of-distribution split. SynRBL produces many balanced and chemically plausible completions, but with lower accuracy on the benchmark test splits. Across all methods, performance degrades with increasing incompleteness. We observe a substantial drop when evaluating on reactions outside the benchmark (full uncurated USPTO), highlighting the gap between benchmark performance and practical robustness and motivating future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。