药物敏感性预测停滞源于数据分布问题,而非模型能力不足。
Training distribution determines the ceiling of drug-blind cancer sensitivity prediction
- 用药物特异性相关性替代全局相关性,更准确评估模型性能
- 仅用细胞特征已能逼近最优表现,药物特征贡献微小
- 按作用机制分组训练可显著提升靶向药预测效果
精准肿瘤学需要根据肿瘤分子特征预测药物抑制效果,但药物盲预测的性能已趋于饱和,尽管药物表示越来越复杂。本文指出,这种停滞是指标偏差所致,而非表示瓶颈。标准基准全局皮尔逊相关系数(Pearson r)受药物间效力差异主导,而一个仅依赖药物均值的简单模型即可捕获此信号。采用每种药物的皮尔逊相关系数(per-drug Pearson r)可隔离药物内部细胞排序性能,结果显示,在四个独立数据集上,任何药物编码均未优于仅使用细胞特征的表现。通过控制实验,将作用机制(MoA)作为特征或训练分布约束,发现将MoA用于分层训练可显著提升靶向激酶抑制剂的预测性能,因泛癌联合训练会抑制通路特异性敏感信号。基于机制分层训练与早期响应匹配的策略,可恢复药物盲预测中的主要增益来源。
原文摘要 · Abstract (English)
Precision oncology requires predicting which drugs will suppress a specific tumor from its molecular profile, but drug-blind sensitivity prediction has plateaued despite increasingly complex drug representations. Here we show that this stagnation reflects a metric artifact rather than a representational bottleneck. The standard benchmark, global Pearson r, is dominated by between-drug potency differences that a trivial drug-mean predictor captures without any cell-specific learning. Per-drug Pearson r, which isolates within-drug cell ranking, reveals that no drug encoding improves over cell-only features across four independent datasets. A controlled experiment channeling mechanism-of-action identity as either a drug feature or a training-distribution constraint identifies the cause. Supplying MoA as a feature yields negligible benefit, whereas using it to stratify training raises per-drug r substantially for targeted kinase inhibitors, because pan-cancer co-training suppresses pathway-specific sensitivity signals. Mechanism-stratified training and response matching from pilot observations provide two deployable strategies that together recover the principal sources of predictive gain in drug-blind sensitivity prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。