用大规模预训练提升核酸-小分子对接精度,5秒生成100种结合构象
An accurate nucleic acid-small molecule docking framework via geometric deep learning with large-scale pretraining
- 融合物理引导的合成数据预训练与实验结构微调
- 在125个测试复合物中成功率达56%(RMSD<2.0Å)
- 适合需快速筛选核酸靶点药物的研究者
核酸作为药物靶点日益受到重视,但小分子与核酸对接的准确性和效率仍面临挑战。传统物理方法精度有限,深度学习受限于实验解析复合物数据稀缺。本文提出NucleoDock框架,通过百万级生成复合物进行物理引导的大规模预训练,并在精选实验共晶结构上微调。模型结合序列与结构信息的核苷酸表示及原子级三维特征,捕捉生物背景与结合位点几何。采用基于混合密度网络的几何评分头,建模条件化相互作用距离分布以排序结合构象。在外部基准125个核酸-配体复合物上,顶1成功率达56%(RMSD截断2.0Å),优于rDock的29%;每复合物约5秒生成100个构象。在ROBIN基准的回顾性虚拟筛选中也显示更好早期富集效果。NucleoDock推动了蛋白与核酸导向计算药物发现之间的方法鸿沟弥合。
原文摘要 · Abstract (English)
Nucleic acids are increasingly recognized as therapeutic targets beyond conventional protein-centered drug discovery, yet accurate and efficient docking of small molecules to nucleic acid structures remains challenging. Physics-based docking methods often show limited accuracy and efficiency, whereas deep learning approaches are constrained by the scarcity of experimentally resolved nucleic acid-ligand complexes. Here, we present NucleoDock, a deep learning framework for nucleic acid-small molecule docking. To address data scarcity, NucleoDock combines physics-guided large-scale pretraining on millions of docking-generated synthetic complexes with fine-tuning on curated experimental co-crystal structures. It further integrates sequence- and structure-informed nucleotide representations with atomistic three-dimensional features to capture both biological context and binding-site geometry. A mixture density network-based geometric scoring head is used to model conditional interaction-distance distributions for pose ranking. On an external benchmark of 125 nucleic acid-ligand complexes, NucleoDock achieved a top-1 success rate of 56 percent at an RMSD cutoff of 2.0 Angstrom, outperforming rDock with 29 percent, while generating 100 poses in approximately 5 seconds per complex. Retrospective virtual screening on the ROBIN benchmark further showed improved early enrichment. NucleoDock represents a step toward bridging the methodological gap between protein- and nucleic acid-directed computational drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。