arXiv:2608.07609stat.APcs.AI2026-08

用SSMD改进药物筛选中单次实验的命中识别效果

Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays

  • 基于SSMD构建机器学习评估指标,适配单次重复数据
  • 在2.2万次单重复实验中验证,结果与传统指标一致
  • 适合低复制数的早期药物筛选场景

高通量筛选(HTS)是新药发现早期的关键环节,但通常因每种化合物仅有一个重复而面临严重数据稀疏问题。传统机器学习评估指标如灵敏度、特异性和受试者工作特征曲线下面积(AUROC)难以准确估计。本文提出一种基于严格标准化均值差(SSMD)的模型化框架,在等方差正态假设下,推导出SSMD与最优约登灵敏度、预设特异性下的灵敏度之间的闭式关系,可从非中心t分布获得精确估计和置信区间,即使在单次重复设计下也适用。与随样本量趋于1的统计功效不同,基于SSMD的灵敏度收敛于反映真实组间差异的有限值,更具实际意义。在包含约2.2万次单重复测量的丙型肝炎病毒siRNA初筛中,验证了基于SSMD、AUROC和灵敏度的阈值所得命中集具有一致性与可解释性。该研究连接了经典HTS统计与机器学习评估理论,为极低复制数筛选流程提供了统计严谨且可复现的性能评估方法。

原文摘要 · Abstract (English)

High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-learning performance metrics, such as sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC), difficult to estimate empirically because they require adequately sized labeled samples. Here, we introduce a model-based framework that derives these classification metrics from the strictly standardized mean difference (SSMD), a well-established HTS effect-size parameter. Under a Gaussian equal-variance assumption, we derive closed-form relationships linking SSMD to Youden-optimal sensitivity and specificity, and sensitivity at a preset specificity, yielding explicit estimators and exact confidence intervals from the noncentral t-distribution, even under single-replicate designs. Unlike classical statistical power, which approaches 1 as sample size grows regardless of how small the true non-zero difference between group means is, the SSMD-derived sensitivity converges to a finite population value that reflects the true degree of separation between two groups, making it a more meaningful and stable performance measure for hit selection. We demonstrate the utility of this framework in a hepatitis C virus primary siRNA screen comprising approximately 22,000 single-replicate measurements, showing that SSMD, AUROC, and sensitivity-based thresholds yield equivalent and interpretable hit sets. This work bridges classical HTS statistics and machine-learning evaluation theory, providing a statistically principled, reproducible way to estimate classification performance in ultra-low-replication screening workflows.

药物筛选机器学习统计评估单重复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。