arXiv:2410.14946cs.LGcs.AI2024-10被引 1

解决DNA编码文库筛选中的读数噪声问题,提升结合亲和力预测准确率。

DEL-Ranking: Ranking-Correction Denoising Framework for Elucidating Molecular Affinities in DNA-Encoded Libraries

  • 引入排序损失修正读数相对大小关系,学习决定活性的因果特征。
  • 通过自训练与一致性损失迭代优化,提升模型对结合活性的预测一致性。
  • 构建首个包含多维分子表征的新型数据集,支持零样本泛化应用。

DNA编码文库(DEL)筛选通过读数实现对蛋白质-配体相互作用的大规模快速探索,但非特异性相互作用导致的读数噪声会误导筛选过程。本文提出DEL-Ranking,一种分布校正去噪框架,包含两项关键创新:(1)新颖的排序损失,用于修正读数间的相对大小关系,使模型能够学习决定活性水平的因果特征;(2)采用自训练与一致性损失的迭代算法,增强活性标签与读数预测之间的模型一致性。此外,我们构建了三个新的DEL筛选数据集,首次全面涵盖多维分子表示、蛋白-配体富集值及活性标签,缓解人工智能驱动的DEL研究中的数据稀缺问题。在多种DEL数据集上的严格评估表明,该模型在多个相关性指标上表现优异,显著提升结合亲和力预测精度,并展现出跨不同蛋白靶标的零样本泛化能力,成功识别出影响化合物结合亲和力的关键结构基序。本工作推动了DEL筛选分析的发展,为未来研究提供重要资源。

原文摘要 · Abstract (English)

DNA-encoded library (DEL) screening has revolutionized the detection of protein-ligand interactions through read counts, enabling rapid exploration of vast chemical spaces. However, noise in read counts, stemming from nonspecific interactions, can mislead this exploration process. We present DEL-Ranking, a novel distribution-correction denoising framework that addresses these challenges. Our approach introduces two key innovations: (1) a novel ranking loss that rectifies relative magnitude relationships between read counts, enabling the learning of causal features determining activity levels, and (2) an iterative algorithm employing self-training and consistency loss to establish model coherence between activity label and read count predictions. Furthermore, we contribute three new DEL screening datasets, the first to comprehensively include multi-dimensional molecular representations, protein-ligand enrichment values, and their activity labels. These datasets mitigate data scarcity issues in AI-driven DEL screening research. Rigorous evaluation on diverse DEL datasets demonstrates DEL-Ranking's superior performance across multiple correlation metrics, with significant improvements in binding affinity prediction accuracy. Our model exhibits zero-shot generalization ability across different protein targets and successfully identifies potential motifs determining compound binding affinity. This work advances DEL screening analysis and provides valuable resources for future research in this area.

DNA编码文库结合亲和力去噪框架零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。