融合三维几何、分子语法和物理化学特征,提升分子属性预测精度
Multimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES Embeddings
- 构建三模态融合网络,整合3D空间、拓扑结构与宏观描述符
- 在QM9数据集上达到0.0207 eV MAE,比基线降低20.6%误差
- 参数少于100万,适合高通量虚拟筛选的高效替代方案
分子性质预测常依赖孤立的数据模态,连续3D图神经网络(GNN)难以高效捕捉长程拓扑依赖和精确的宏观启发式规律。本文提出一种参数高效的三分支模块化融合神经网络,融合三种正交模态:3D空间几何(SchNet)、离散拓扑语法(通过ChemBERTa获取的SMILES)以及显式的宏观物化描述符(深度与交叉网络)。通过跳过标准标量读出并采用共享的晚期融合架构,该框架建立了一个数学严谨的多模态潜在空间,有效解决局部消息传递带来的算术与过度平滑问题。在QM9基准上评估,目标为0K下的原子化能($U_0^{\mathrm{atom}}$)热力学性质。通过系统的组合消融与潜在瓶颈优化($d_e=64$),三模态框架在验证集上实现0.0207 eV的平均绝对误差(MAE)。模型参数少于一百万,在突破化学精度阈值的同时,相较严格控制的几何基线实现20.6%的误差降低。结果表明,整合正交的宏观与拓扑数据流可提供$Ω(1)$物理捷径,为暴力参数扩展提供高效替代方案,适用于高通量虚拟筛选(HTVS)流程的稳健代理模型。
原文摘要 · Abstract (English)
Molecular property prediction often relies on isolated data modalities, where continuous 3D graph neural networks (GNNs) struggle to efficiently capture long-range topological dependencies and exact macroscopic heuristics. In this work, we introduce a parameter-efficient Tri-Branch Modular Fusion Neural Network that synthesizes three orthogonal modalities: 3D spatial geometry (SchNet), discrete topological grammar (SMILES via ChemBERTa), and explicit macroscopic physicochemical descriptors (Deep & Cross Network). By bypassing standard scalar readouts and employing a shared late-fusion architecture, the framework establishes a mathematically rigorous multimodal latent space that effectively resolves the arithmetic and oversmoothing limitations of local message passing. We evaluate the proposed architecture on the QM9 benchmark, targeting the extensive thermodynamic property of atomization energy at 0 K ($U_0^{\mathrm{atom}}$). Through systematic combinatorial ablation and latent bottleneck optimization ($d_e=64$), the tri-modal framework achieves a validation Mean Absolute Error (MAE) of 0.0207 eV. Operating with fewer than one million parameters, this architecture decisively surpasses the sub-chemical accuracy threshold and yields a substantial 20.6% error reduction over a strictly controlled geometric baseline. Ultimately, our findings demonstrate that integrating orthogonal macroscopic and topological data streams provides a synergistic, $\mathcal{O}(1)$ physical shortcut. This multimodal alignment offers a highly efficient alternative to brute-force parameter scaling, establishing a robust surrogate model for high-throughput virtual screening (HTVS) pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。