通过子结构对齐提升分子与文本描述的匹配精度
Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignment
- 基于分子子结构与化学短语的细粒度对齐信号增强
- 在多个分子基准上超越现有模型,显著提升匹配效果
- 适合药物发现、化学信息理解等需要精准匹配的场景
分子与文本表示学习因提升化学信息理解潜力而受到关注。然而,现有模型难以捕捉分子与其描述之间的细微差异,因其缺乏对分子子结构与化学短语之间细粒度对齐的学习能力。为此,我们提出MolBridge,一种基于子结构感知对齐的新型分子-文本学习框架。具体地,我们通过分子子结构和化学短语生成额外的对齐信号,增强原始分子-描述配对。为有效利用这些丰富对齐信息,MolBridge采用子结构感知对比学习,并引入自精炼机制过滤噪声对齐信号。实验表明,MolBridge能有效捕获细粒度对应关系,在多种分子基准上优于现有最优模型,凸显子结构感知对齐在分子-文本学习中的重要性。
原文摘要 · Abstract (English)
Molecule and text representation learning has gained increasing interest due to its potential for enhancing the understanding of chemical information. However, existing models often struggle to capture subtle differences between molecules and their descriptions, as they lack the ability to learn fine-grained alignments between molecular substructures and chemical phrases. To address this limitation, we introduce MolBridge, a novel molecule-text learning framework based on substructure-aware alignments. Specifically, we augment the original molecule-description pairs with additional alignment signals derived from molecular substructures and chemical phrases. To effectively learn from these enriched alignments, MolBridge employs substructure-aware contrastive learning, coupled with a self-refinement mechanism that filters out noisy alignment signals. Experimental results show that MolBridge effectively captures fine-grained correspondences and outperforms state-of-the-art baselines on a wide range of molecular benchmarks, highlighting the significance of substructure-aware alignment in molecule-text learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。