选不同评价指标,谁赢谁输完全反转,揭示药物响应预测的评测陷阱。
The Metric Picks the Winner: Evaluation Choice Flips Model Rankings for Drug-Response Prediction in Unseen Chemistry

- 用融合化学嵌入与检索特征的模型预测基因表达残差。
- 在真实评分下,深度模型显著优于线性指纹基线(-0.012 wMSE)。
- 首次在真实未见化学结构上验证了评测指标对模型排名的决定性影响。
预测细胞对未见过药物的转录组响应是计算细胞生物学的核心难题。在THP-1细胞DRUG-seq数据集上,基于活性化合物加权均方误差(wMSE)的竞赛评分标准下,采用Bemis-Murcko骨架划分时,模型排名随评价指标变化而反转:使用基因层面逆方差代理指标时,基于摩根指纹的正则化线性回归表现最佳;但采用竞赛真实评分(按基因-化合物对加权梅吉亚权重)时,深度模型胜出,本研究提出的融合解码器显著优于线性基线(-0.012 wMSE,配对置换检验p < 10^-4),而代理指标下的优胜者反而成为最差的化学感知预测器。结果表明,评价指标的选择直接决定了模型胜负——这是首次在真实未见药物化学结构上验证该现象。我们公开可复现的流水线,支持向真实1064×12,995测试网格提交有效结果。
原文摘要 · Abstract (English)
Predicting how a cell's transcriptome responds to a drug it has never seen is a core, hard problem in computational cell biology: recent benchmarks show complex models often fail to beat trivial baselines once test compounds are held out by chemistry. We study one cell line and assay, THP-1 cells profiled by DRUG-seq, scored by the active-compound weighted MSE(wMSE) of the VCPI prediction contest. We propose a staged approach: dumb baselines (untreated control and mean training-compound response) that the field keeps failing to beat; non-parametric retrieval (a Tanimoto-weighted average of a held-out compound's nearest training compounds); and a fusion stage combining a frozen chemistry embedding with retrieval-support features to predict the residual over the mean, with an uncertainty head and gene programs. On the released VCPI THP-1 drug-seq data (14,026 training compounds), under a Bemis-Murcko scaffold split, the model ranking inverts depending on the metric. Under an inverse-variance per-gene proxy, a regularized linear regression on Morgan fingerprints appears to win over the deep models, retrieval, and ChemBERTa -- the textbook "simple baselines win" result. But under the contest's true active-set metric (per-(gene, compound) Mejia weights, validated against the official scorer; mean baseline 0.535 vs the organizers' 0.507 reference), that reverses: the deep models win, our fusion decoder significantly beats the linear fingerprint baseline (-0.012 wMSE, paired bootstrap p < 10^-4), and the proxy's winner becomes the worst chemistry-aware predictor. Picking the metric picks the winner -- to our knowledge the first demonstration on real held-out drug chemistry of the metric-calibration effect established largely on genetic perturbation. We release a reproducible pipeline wired to the official scorer that emits a valid submission over the real 1064 x 12,995 grid.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。