用化学构象变化机制提升分子关系学习的稳定性
Representational Alignment with Chemical Induced Fit for Molecular Relational Learning
- 引入化学诱导契合原理,动态对齐分子子结构表示
- 在九个数据集上优于现有模型,尤其在骨架迁移时表现更稳
- 适合分子属性预测与药物设计场景,提升模型泛化能力
分子关系学习(MRL)广泛用于自然科学研究中,通过提取结构特征预测分子对之间的关系。子结构对之间的表示相似性决定了分子结合位点的功能兼容性。然而,仅依赖注意力机制对齐子结构表示缺乏化学知识指导,导致模型在化学空间(如官能团、骨架)偏移数据上性能不稳定。本文提出代表式对齐与化学诱导契合(ReAlignFit),通过引入基于化学诱导契合的归纳偏置,动态对齐MRL中的子结构表示。在归纳过程中,设计基于子结构边重建的偏差校正函数,模拟化学构象变化(子结构的动态组合),实现表示对齐。同时,在拟合过程中集成子图信息瓶颈,优化高化学功能兼容性的子结构对,生成分子嵌入。在九个数据集上的实验表明,ReAlignFit在两项任务中超越当前最优模型,并显著提升模型在规则偏移和骨架偏移数据分布下的稳定性。
原文摘要 · Abstract (English)
Molecular Relational Learning (MRL) is widely applied in natural sciences to predict relationships between molecular pairs by extracting structural features. The representational similarity between substructure pairs determines the functional compatibility of molecular binding sites. Nevertheless, aligning substructure representations by attention mechanisms lacks guidance from chemical knowledge, resulting in unstable model performance in chemical space (\textit{e.g.}, functional group, scaffold) shifted data. With theoretical justification, we propose the \textbf{Re}presentational \textbf{Align}ment with Chemical Induced \textbf{Fit} (ReAlignFit) to enhance the stability of MRL. ReAlignFit dynamically aligns substructure representation in MRL by introducing chemical Induced Fit-based inductive bias. In the induction process, we design the Bias Correction Function based on substructure edge reconstruction to align representations between substructure pairs by simulating chemical conformational changes (dynamic combination of substructures). ReAlignFit further integrates the Subgraph Information Bottleneck during fit process to refine and optimize substructure pairs exhibiting high chemical functional compatibility, leveraging them to generate molecular embeddings. Experimental results on nine datasets demonstrate that ReAlignFit outperforms state-of-the-art models in two tasks and significantly enhances model's stability in both rule-shifted and scaffold-shifted data distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。