arXiv:2608.11444q-bio.QMcs.LG2026-08

扩充抗癌药效预测数据集,提升模型对新化合物的泛化能力。

Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

论文配图:Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling
图 1 · 摘自论文原文
  • 整合PharmacoDB等数据源,构建超大规模药理基因组数据集。
  • 新数据集包含超5万种化合物,显著提升化学空间多样性。
  • 模型在新化合物和隔离数据上表现更优,适合药物发现研究。

药物反应预测(DRP)是药物基因组学中的活跃研究方向,有望加速抗癌药物的有效性筛选。然而,现有模型性能受限于数据规模不足及癌症与化学空间覆盖有限。同时,不一致的评估方法阻碍了模型间的可靠比较。标准化框架如IMPROVE项目提供了统一的数据模式和评估协议,但提升模型泛化能力仍需更大更丰富的训练数据。本文通过大规模整合PharmacoDB及其他小规模数据源,显著扩展了IMPROVE基准数据集,涵盖数百万药物反应测量值,拓展了多组学覆盖范围,并新增超过50,000种化合物,大幅提高化学多样性。为评估新数据集效果,我们使用原版与扩展版数据分别训练模型,在相同测试集和多种评估策略(包括药物盲、癌症盲、分离数据划分)下进行对比。结果表明,尽管癌症盲性能与原基准相当,但基于扩展数据训练的模型在药物盲和分离数据设置中均表现出持续提升,表明其对未见化合物具有更强泛化能力。该扩展数据集可作为社区资源,为开发辅助新型抗癌药物发现的DRP模型提供更坚实基础。

原文摘要 · Abstract (English)

Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer and chemical spaces. In addition, inconsistent benchmarking practices hinder reliable comparison across models. Standardized frameworks, such as the Innovative Methodologies and New Data for Predictive Oncology Model Evaluation (IMPROVE) project, provide unified data schemas and evaluation protocols for consistent benchmarking, but improving model generalizability requires larger and more diverse training data. In this work, we substantially expand the IMPROVE benchmark through large-scale integration of pharmacogenomic data, primarily from PharmacoDB, together with additional smaller data sources. The expanded resource includes millions of drug response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds. To evaluate the impact of the new dataset compared to the original IMPROVE benchmark dataset, we trained DRP models using the two datasets and assess their prediction performance using a common test set and several evaluation strategies, including drug-blind, cancer-blind, and disjoint data splits. While cancer-blind performance remained comparable to the original benchmark, models trained on the expanded dataset showed consistent improvements in drug-blind and disjoint settings, indicating enhanced generalization to previously unseen compounds. These results position the expanded dataset as a community resource that provides a richer foundation for developing DRP models intended to aid in the discovery of novel anticancer drugs.

药物反应预测抗癌药物大规模数据多组学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。