arXiv:2410.23715cs.IR2024-10被引 5

提升文本与分子跨模态检索精度,通过共享特征和二阶相似性对齐。

Towards Cross-Modal Text-Molecule Retrieval with Better Modality Alignment

  • 引入可学习记忆向量的特征投影器,更好提取文本与分子的共享特征。
  • 基于四类相似度分布的二阶相似性损失,增强跨模态对齐效果。
  • 在多个数据集上达到当前最优性能,提升幅度达6.4%。

跨模态文本-分子检索模型旨在学习文本与分子模态的共享特征空间,以实现精准的相似性计算,从而加速药物设计中具有特定性质和活性的分子筛选。然而,以往方法存在两大缺陷:一是未能充分捕捉文本序列与分子图之间的显著差异所导致的模态共享特征;二是主要依赖对比学习和对抗训练进行跨模态对齐,两者均聚焦于一阶相似性,忽视了能捕获嵌入空间更丰富结构信息的二阶相似性。为此,本文提出一种新型跨模态文本-分子检索模型,包含两方面改进:首先,在两个模态专用编码器基础上,堆叠基于记忆库的特征投影器,其中包含可学习的记忆向量,以更好地提取模态共享特征;更重要的是,在训练过程中,为每个实例计算四类相似度分布(文本-文本、文本-分子、分子-分子、分子-文本),并通过最小化这些相似度分布间的距离(即二阶相似性损失)来增强跨模态对齐。实验结果与分析充分证明了该模型的有效性。尤其在多个基准数据集上,模型表现超越此前最优结果6.4%,达到当前最佳水平。

原文摘要 · Abstract (English)

Cross-modal text-molecule retrieval model aims to learn a shared feature space of the text and molecule modalities for accurate similarity calculation, which facilitates the rapid screening of molecules with specific properties and activities in drug design. However, previous works have two main defects. First, they are inadequate in capturing modality-shared features considering the significant gap between text sequences and molecule graphs. Second, they mainly rely on contrastive learning and adversarial training for cross-modality alignment, both of which mainly focus on the first-order similarity, ignoring the second-order similarity that can capture more structural information in the embedding space. To address these issues, we propose a novel cross-modal text-molecule retrieval model with two-fold improvements. Specifically, on the top of two modality-specific encoders, we stack a memory bank based feature projector that contain learnable memory vectors to extract modality-shared features better. More importantly, during the model training, we calculate four kinds of similarity distributions (text-to-text, text-to-molecule, molecule-to-molecule, and molecule-to-text similarity distributions) for each instance, and then minimize the distance between these similarity distributions (namely second-order similarity losses) to enhance cross-modal alignment. Experimental results and analysis strongly demonstrate the effectiveness of our model. Particularly, our model achieves SOTA performance, outperforming the previously-reported best result by 6.4%.

跨模态检索分子生成特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。