用大模型和多模态数据预测药物靶点结合,提升新药研发效率。
LLM$^3$-DTI: A Large Language Model and Multi-modal data co-powered framework for Drug-Target Interaction prediction
- 融合药物/靶点文本的领域大模型与多模态特征
- 双交叉注意力与TSFusion模块提升特征对齐与融合
- 在多个数据集上超越现有方法,适合药物研发人员
药物-靶点相互作用(DTI)预测对新药研发和药物重定位具有重要意义。随着海量数据积累,数据驱动方法被广泛用于预测DTI,显著降低研发成本。本文提出一种基于大语言模型与多模态数据协同的药物-靶点相互作用预测框架——LLM$^3$-DTI。该框架通过领域专用大模型编码药物与靶点的文本语义嵌入,构建多模态数据表示。为有效对齐与融合多模态嵌入,提出双交叉注意力机制与TSFusion模块,并通过输出网络完成DTI任务。实验结果表明,LLM$^3$-DTI能高效识别已验证的DTI,在多种场景下均优于对比模型。代码与数据已开源:https://github.com/chaser-gua/LLM3DTI。
原文摘要 · Abstract (English)
Drug-target interaction (DTI) prediction is of great significance for drug discovery and drug repurposing. With the accumulation of a large volume of valuable data, data-driven methods have been increasingly harnessed to predict DTIs, reducing costs across various dimensions. Therefore, this paper proposes a $\textbf{L}$arge $\textbf{L}$anguage $\textbf{M}$odel and $\textbf{M}$ulti-$\textbf{M}$odel data co-powered $\textbf{D}$rug $\textbf{T}$arget $\textbf{I}$nteraction prediction framework, named LLM$^3$-DTI. LLM$^3$-DTI constructs multi-modal data embedding to enhance DTI prediction performance. In this framework, the text semantic embeddings of drugs and targets are encoded by a domain-specific LLM. To effectively align and fuse multi-modal embedding. We propose the dual cross-attention mechanism and the TSFusion module. Finally, these multi-modal data are utilized for the DTI task through an output network. The experimental results indicate that LLM$^3$-DTI can proficiently identify validated DTIs, surpassing the performance of the models employed for comparison across diverse scenarios. Consequently, LLM$^3$-DTI is adept at fulfilling the task of DTI prediction with excellence. The data and code are available at https://github.com/chaser-gua/LLM3DTI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。