用视觉模型预测反应产率,比传统方法更准且可解释。
Generic Vision and Cross-Attention for Reaction Yield Prediction

- 融合分子二维结构与物理数据,用跨注意力机制协同学习。
- 测试集均方根误差达5.27%,优于纯量子描述符方法。
- 能识别关键位阻瓶颈,适合化学研发与深度学习交叉研究者。
传统反应产率预测受限于缺乏空间信息的1D量子描述符。为此,提出一种双模态视觉-跨注意力架构,融合表格型物化数据与2D分子拓扑。显著发现,仅用通用视觉主干处理简单2D骨架即可超越纯量子基线。通过协同双模态,最优跨注意力框架在测试集上达到5.27%的均方根误差(RMSE),优于传统方法。机制分析显示,网络实现基于描述符引导的空间查询,将宏观位阻识别任务有效转移至视觉路径。同时,网络动态学习化学层次结构,高度关注关键位阻位点(如芳基卤化物)。残差跳跃连接则保护非空间电子参数免受融合过程中的破坏性衰减。整体提供了一种可扩展、高可解释的深度视觉学习增强物理化学的范式。
原文摘要 · Abstract (English)
Traditional reaction yield prediction is constrained by 1D quantum descriptors that lack explicit spatial information. To address this gap, a dual-modal Vision Cross-Attention architecture is proposed, fusing tabular physical-organic data with 2D molecular topologies. Notably, it is demonstrated that a generic computer vision backbone processing simple 2D skeletal structures independently outperforms purely quantum-based baselines. By synergizing both modalities, superior predictive accuracy compared to traditional methodologies is achieved by the optimal cross-attention framework (Test RMSE = 5.27%). Through mechanistic probing, active, descriptor-guided spatial querying is observed, effectively offloading macroscopic steric identification to the visual pathway. Furthermore, a dynamic chemical hierarchy is learned by the network to heavily prioritize critical steric bottlenecks, such as the aryl halide. Concurrently, residual skip connections are utilized to protect non-spatial electronic parameters from destructive attenuation during fusion. Collectively, a scalable and highly interpretable blueprint is provided for augmenting physical chemistry with deep visual learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。