arXiv:2509.16273cs.LGcs.AI2025-09

通过子图驱动的动态传播,提升小样本药物筛选准确率

SubDyve: Subgraph-Driven Dynamic Propagation for Virtual Screening Enhancement Controlling False Positive

  • 构建子图感知相似网络,动态传播已知活性分子信号
  • 在仅少量已知活性分子时,仍保持高精度,最高提升34.0分
  • 适合低标签场景下的新药筛选,尤其适用于数据稀缺任务

虚拟筛选(VS)旨在从庞大化学库中识别具有生物活性的化合物,但在仅有少量已知活性化合物的低标签环境下仍具挑战。现有方法多依赖通用分子指纹,忽视对生物活性关键的类别区分性子结构,且将分子独立处理,限制了在低标签场景下的效果。我们提出SubDyve,一种基于图网络的虚拟筛选框架,构建子图感知相似网络,并从少量已知活性化合物出发传播活性信号。当活性化合物极少时,SubDyve采用迭代种子精炼策略,基于局部假发现率逐步提升候选分子。该策略在扩大种子集的同时,有效控制拓扑偏差和过度扩展带来的假阳性。我们在十个人类靶点(DUD-E)的零样本条件下,以及包含一千万化合物的ZINC数据集上的CDK7靶点上评估了SubDyve,结果表明其持续优于现有的指纹或嵌入式方法,在BEDROC指标上最高提升34.0,在EF1%指标上最高提升24.6。

原文摘要 · Abstract (English)

Virtual screening (VS) aims to identify bioactive compounds from vast chemical libraries, but remains difficult in low-label regimes where only a few actives are known. Existing methods largely rely on general-purpose molecular fingerprints and overlook class-discriminative substructures critical to bioactivity. Moreover, they consider molecules independently, limiting effectiveness in low-label regimes. We introduce SubDyve, a network-based VS framework that constructs a subgraph-aware similarity network and propagates activity signals from a small known actives. When few active compounds are available, SubDyve performs iterative seed refinement, incrementally promoting new candidates based on local false discovery rate. This strategy expands the seed set with promising candidates while controlling false positives from topological bias and overexpansion. We evaluate SubDyve on ten DUD-E targets under zero-shot conditions and on the CDK7 target with a 10-million-compound ZINC dataset. SubDyve consistently outperforms existing fingerprint or embedding-based approaches, achieving margins of up to +34.0 on the BEDROC and +24.6 on the EF1% metric.

虚拟筛选子图学习低标签学习药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。