arXiv:2608.00776cs.LGphysics.chem-ph2026-08

用视觉模型预测反应产率,比传统方法更准且可解释。

Generic Vision and Cross-Attention for Reaction Yield Prediction

论文配图:Generic Vision and Cross-Attention for Reaction Yield Prediction
图 1 · 摘自论文原文
  • 融合分子二维结构与物理数据,用跨注意力机制协同学习。
  • 测试集均方根误差达5.27%,优于纯量子描述符方法。
  • 能识别关键位阻瓶颈,适合化学研发与深度学习交叉研究者。

传统反应产率预测受限于缺乏空间信息的1D量子描述符。为此,提出一种双模态视觉-跨注意力架构,融合表格型物化数据与2D分子拓扑。显著发现,仅用通用视觉主干处理简单2D骨架即可超越纯量子基线。通过协同双模态,最优跨注意力框架在测试集上达到5.27%的均方根误差(RMSE),优于传统方法。机制分析显示,网络实现基于描述符引导的空间查询,将宏观位阻识别任务有效转移至视觉路径。同时,网络动态学习化学层次结构,高度关注关键位阻位点(如芳基卤化物)。残差跳跃连接则保护非空间电子参数免受融合过程中的破坏性衰减。整体提供了一种可扩展、高可解释的深度视觉学习增强物理化学的范式。

原文摘要 · Abstract (English)

Traditional reaction yield prediction is constrained by 1D quantum descriptors that lack explicit spatial information. To address this gap, a dual-modal Vision Cross-Attention architecture is proposed, fusing tabular physical-organic data with 2D molecular topologies. Notably, it is demonstrated that a generic computer vision backbone processing simple 2D skeletal structures independently outperforms purely quantum-based baselines. By synergizing both modalities, superior predictive accuracy compared to traditional methodologies is achieved by the optimal cross-attention framework (Test RMSE = 5.27%). Through mechanistic probing, active, descriptor-guided spatial querying is observed, effectively offloading macroscopic steric identification to the visual pathway. Furthermore, a dynamic chemical hierarchy is learned by the network to heavily prioritize critical steric bottlenecks, such as the aryl halide. Concurrently, residual skip connections are utilized to protect non-spatial electronic parameters from destructive attenuation during fusion. Collectively, a scalable and highly interpretable blueprint is provided for augmenting physical chemistry with deep visual learning.

反应产率预测视觉模型跨注意力可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。