arXiv:2409.14106cs.AI2024-09KDD被引 9

让分子图与文本在原子级细节上对齐,提升模型对新分子的预测能力。

Advancing Molecular Graph-Text Pre-training via Fine-grained Alignment

  • 通过对比学习和掩码多模态建模,同时捕捉分子整体与关键片段的特征。
  • 在文本驱动的分子编辑任务中性能提升最高达230%。
  • 适合需要精细理解分子结构的药物研发与催化剂设计场景。

理解分子结构及其相关知识对科研至关重要。近期研究将分子图与其文字描述结合以增强分子表征学习,但主要关注整个分子图,忽略了频繁出现的、决定分子性质的关键子图(即基元)。缺乏细粒度知识导致模型难以泛化到未见分子及需要基元级洞察的任务。为此,我们提出 FineMolTex,一种细粒度分子图-文本预训练框架,联合学习粗粒度分子级知识与细粒度基元级知识。具体包括:基于对比对齐的粗粒度匹配任务,以及基于重要性选择的掩码多模态建模任务,后者需预测被遮蔽的基元和词语标签。通过融合双模态信息,FineMolTex 能够实现基元与文本间的细粒度对齐。我们在三个下游任务上进行了大量实验,文本驱动的分子编辑任务性能最高提升230%。案例研究还表明,FineMolTex 成功捕捉了细粒度知识,为药物发现与催化剂设计提供潜在价值。

原文摘要 · Abstract (English)

Understanding molecular structure and related knowledge is crucial for scientific research. Recent studies integrate molecular graphs with their textual descriptions to enhance molecular representation learning. However, they focus on the whole molecular graph and neglect frequently occurring subgraphs, known as motifs, which are essential for determining molecular properties. Without such fine-grained knowledge, these models struggle to generalize to unseen molecules and tasks that require motif-level insights. To bridge this gap, we propose FineMolTex, a novel Fine-grained Molecular graph-Text pre-training framework to jointly learn coarse-grained molecule-level knowledge and fine-grained motif-level knowledge. Specifically, FineMolTex consists of two pre-training tasks: a contrastive alignment task for coarse-grained matching and a masked multi-modal modeling task for fine-grained matching. In particular, the latter predicts the labels of masked motifs and words, which are selected based on their importance. By leveraging insights from both modalities, FineMolTex is able to understand the fine-grained matching between motifs and words. Finally, we conduct extensive experiments across three downstream tasks, achieving up to 230% improvement in the text-based molecule editing task. Additionally, our case studies reveal that FineMolTex successfully captures fine-grained knowledge, potentially offering valuable insights for drug discovery and catalyst design.

分子图多模态预训练药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。