arXiv:2410.08938q-bio.QMcs.LG2024-10ICML被引 4

公开首个含结合构象的激酶抑制剂DNA编码库数据集,助力机器学习药物发现。

KinDEL: DNA-Encoded Library Dataset for Kinase Inhibitors

  • 构建8100万化合物的激酶抑制剂DEL数据集,包含2D/3D结构与结合构象。
  • 首次提供基于分子对接的结合姿态数据,支持结构建模与预测。
  • 适合从事药物发现、机器学习或计算化学的研究者使用。

DNA编码库(DEL)是药物发现中变革性的技术,可高通量探索巨大的化学空间。然而,公开可用的DEL数据集稀缺,制约了该领域机器学习方法的发展。为此,我们推出KinDEL,目前最大的公开可访问DEL数据集之一,也是首个包含分子对接结合姿态的数据集。聚焦于两种激酶:丝裂原活化蛋白激酶14(MAPK14)和盘状结构域受体酪氨酸激酶1(DDR1),KinDEL包含8100万种化合物,为计算研究提供丰富资源。此外,我们还提供了全面的生物物理实验验证数据,涵盖在DNA上与离DNA的测量结果,并用于评估多种机器学习方法,包括新型基于结构的概率模型。我们希望这一基准数据集(包含2D和3D结构)能推动基于数据驱动的DEL命中识别机器学习模型的发展。

原文摘要 · Abstract (English)

DNA-Encoded Libraries (DELs) represent a transformative technology in drug discovery, facilitating the high-throughput exploration of vast chemical spaces. Despite their potential, the scarcity of publicly available DEL datasets presents a bottleneck for the advancement of machine learning methodologies in this domain. To address this gap, we introduce KinDEL, one of the largest publicly accessible DEL datasets and the first one that includes binding poses from molecular docking experiments. Focused on two kinases, Mitogen-Activated Protein Kinase 14 (MAPK14) and Discoidin Domain Receptor Tyrosine Kinase 1 (DDR1), KinDEL includes 81 million compounds, offering a rich resource for computational exploration. Additionally, we provide comprehensive biophysical assay validation data, encompassing both on-DNA and off-DNA measurements, which we use to evaluate a suite of machine learning techniques, including novel structure-based probabilistic models. We hope that our benchmark, encompassing both 2D and 3D structures, will help advance the development of machine learning models for data-driven hit identification using DELs.

DNA编码库激酶抑制剂机器学习药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。