通过多尺度融合提升DNA编码库去噪能力,发现高亲和力药物候选分子。
Unlocking Potential Binders: Multimodal Pretraining DEL-Fusion for Denoising DNA-Encoded Libraries
- 构建多模态预训练框架,融合原子、片段与分子级特征
- 在三个数据集上显著提升去噪效果,识别出更多高亲和力分子
- 适合药物发现领域研究人员,尤其关注复杂筛选数据的净化
在药物发现领域,DNA编码库(DEL)筛选技术已成为高效识别高亲和力化合物的方法。然而,复杂生物系统中的非特异性相互作用导致显著噪声。已有神经网络模型基于DEL库训练以提取化合物特征,实现数据去噪。但DEL固有的结构限制(如单体多样性有限)影响了化合物编码器性能,且现有方法仅捕捉单一层次的特征,制约去噪效果。为此,本文提出多模态预训练DEL融合模型(MPDF),通过预训练增强编码器能力,并整合不同尺度的化合物特征。设计对比学习任务,关联化合物表示与其文本描述,提升编码器泛化特征提取能力。提出新型DEL融合框架,综合原子、子分子和分子层级信息,由多个编码器捕获。二者协同使MPDF具备丰富多尺度特征,支持全面下游去噪。在三个DEL数据集上评估,MPDF在验证任务中表现更优,揭示了识别高亲和力分子的新路径,推动DEL在药物发现中的应用。
原文摘要 · Abstract (English)
In the realm of drug discovery, DNA-encoded library (DEL) screening technology has emerged as an efficient method for identifying high-affinity compounds. However, DEL screening faces a significant challenge: noise arising from nonspecific interactions within complex biological systems. Neural networks trained on DEL libraries have been employed to extract compound features, aiming to denoise the data and uncover potential binders to the desired therapeutic target. Nevertheless, the inherent structure of DEL, constrained by the limited diversity of building blocks, impacts the performance of compound encoders. Moreover, existing methods only capture compound features at a single level, further limiting the effectiveness of the denoising strategy. To mitigate these issues, we propose a Multimodal Pretraining DEL-Fusion model (MPDF) that enhances encoder capabilities through pretraining and integrates compound features across various scales. We develop pretraining tasks applying contrastive objectives between different compound representations and their text descriptions, enhancing the compound encoders' ability to acquire generic features. Furthermore, we propose a novel DEL-fusion framework that amalgamates compound information at the atomic, submolecular, and molecular levels, as captured by various compound encoders. The synergy of these innovations equips MPDF with enriched, multi-scale features, enabling comprehensive downstream denoising. Evaluated on three DEL datasets, MPDF demonstrates superior performance in data processing and analysis for validation tasks. Notably, MPDF offers novel insights into identifying high-affinity molecules, paving the way for improved DEL utility in drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。