用最优传输实现文本与分子多粒度对齐,提升化学检索精度
Exploring Optimal Transport-Based Multi-Grained Alignments for Text-Molecule Retrieval
- 引入最优传输机制,实现文本词元与分子片段的多粒度对齐
- 在ChEBI-20和PCdes数据集上超越现有SOTA模型,显著提升检索性能
- 适合从事药物发现、化学信息学的科研人员参考
生物信息学领域进展迅速,跨模态文本-分子检索任务日益重要。该任务旨在根据文本描述准确检索分子结构,通过有效对齐文本与分子,帮助研究人员筛选合适分子候选物。然而,现有方法常忽略分子子结构中的细节。本文提出最优传输驱动的多粒度对齐模型(ORMA),通过文本编码器生成词元级与句级表示,将分子建模为包含原子、基团和分子节点的分层异构图,提取三层次表征。关键创新在于利用最优传输(OT)对齐词元与基团,生成融合多词元-基团匹配的复合表示。同时,采用对比学习在词元-原子、多词元-基团、句-分子三个尺度优化跨模态对齐,最大化正确匹配对的相似性,最小化错误匹配对的相似性。据我们所知,这是首次探索基团与多词元层级对齐的工作。在ChEBI-20和PCdes数据集上的实验表明,ORMA显著优于现有SOTA模型。
原文摘要 · Abstract (English)
The field of bioinformatics has seen significant progress, making the cross-modal text-molecule retrieval task increasingly vital. This task focuses on accurately retrieving molecule structures based on textual descriptions, by effectively aligning textual descriptions and molecules to assist researchers in identifying suitable molecular candidates. However, many existing approaches overlook the details inherent in molecule sub-structures. In this work, we introduce the Optimal TRansport-based Multi-grained Alignments model (ORMA), a novel approach that facilitates multi-grained alignments between textual descriptions and molecules. Our model features a text encoder and a molecule encoder. The text encoder processes textual descriptions to generate both token-level and sentence-level representations, while molecules are modeled as hierarchical heterogeneous graphs, encompassing atom, motif, and molecule nodes to extract representations at these three levels. A key innovation in ORMA is the application of Optimal Transport (OT) to align tokens with motifs, creating multi-token representations that integrate multiple token alignments with their corresponding motifs. Additionally, we employ contrastive learning to refine cross-modal alignments at three distinct scales: token-atom, multitoken-motif, and sentence-molecule, ensuring that the similarities between correctly matched text-molecule pairs are maximized while those of unmatched pairs are minimized. To our knowledge, this is the first attempt to explore alignments at both the motif and multi-token levels. Experimental results on the ChEBI-20 and PCdes datasets demonstrate that ORMA significantly outperforms existing state-of-the-art (SOTA) models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。