建立标准化框架评估药物反应预测模型跨数据集泛化能力
Benchmarking community drug response prediction models: datasets, models, tools, and metrics for cross-dataset generalization analysis
- 构建包含5个公开数据集和6种模型的统一评估框架
- 发现模型在新数据集上性能普遍下降,最高降幅显著
- 指出CTRPv2数据集最适合作为训练源,适合药物研发人员参考
深度学习与机器学习模型在药物反应预测中展现潜力,但其跨数据集泛化能力仍存疑,影响实际应用。由于缺乏标准化评估方法,现有模型比较常基于不一致的数据集与指标,难以真实反映预测能力。本文提出一个基准评估框架,整合五个公开药物筛选数据集、六种标准DRP模型及可扩展的工作流程,用于系统评估跨数据集泛化性。引入一组评估指标,量化绝对性能(如跨数据集预测准确率)与相对性能(如相比原数据集的性能下降幅度),实现更全面的迁移能力分析。结果表明,模型在未见数据集上表现大幅下降,凸显严格泛化评估的重要性。虽部分模型具备较强跨数据集泛化能力,但无一模型在所有数据集上持续领先。此外,我们发现CTRPv2作为训练数据源时,在目标数据集上获得更高泛化得分。通过向社区共享该标准化框架,本研究旨在建立严谨的模型比较基础,推动可用于真实场景的鲁棒性药物反应预测模型发展。
原文摘要 · Abstract (English)
Deep learning (DL) and machine learning (ML) models have shown promise in drug response prediction (DRP), yet their ability to generalize across datasets remains an open question, raising concerns about their real-world applicability. Due to the lack of standardized benchmarking approaches, model evaluations and comparisons often rely on inconsistent datasets and evaluation criteria, making it difficult to assess true predictive capabilities. In this work, we introduce a benchmarking framework for evaluating cross-dataset prediction generalization in DRP models. Our framework incorporates five publicly available drug screening datasets, six standardized DRP models, and a scalable workflow for systematic evaluation. To assess model generalization, we introduce a set of evaluation metrics that quantify both absolute performance (e.g., predictive accuracy across datasets) and relative performance (e.g., performance drop compared to within-dataset results), enabling a more comprehensive assessment of model transferability. Our results reveal substantial performance drops when models are tested on unseen datasets, underscoring the importance of rigorous generalization assessments. While several models demonstrate relatively strong cross-dataset generalization, no single model consistently outperforms across all datasets. Furthermore, we identify CTRPv2 as the most effective source dataset for training, yielding higher generalization scores across target datasets. By sharing this standardized evaluation framework with the community, our study aims to establish a rigorous foundation for model comparison, and accelerate the development of robust DRP models for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。