新评估方法解决单细胞扰动模型模式崩溃问题
Diversity by Design: Addressing Mode Collapse Improves scRNA-seq Perturbation Modeling on Well-Calibrated Metrics
- 设计了关注差异表达基因的加权误差指标
- 改进后真实预测模型得分显著高于均值基线
- 适合做单细胞扰动建模与评估的研究者参考
近期基准测试显示,单细胞扰动响应模型常被简单预测数据均值的方法超越。我们发现这是由于评估指标缺陷:以对照组为参考的差值和未加权误差指标,在对照组有偏差或生物信号稀疏时,会奖励模式崩溃。大规模仿真及两个真实扰动数据集分析表明,共享参考偏移而非真实生物变化导致高分表现。本文提出针对所有扰动的差异表达基因(DEG)感知指标:加权均方误差(WMSE)和加权差值决定系数($R^{2}_{w}(Δ)$),能高灵敏度检测微弱信号。同时引入正负性能基线校准指标。改进后,均值基线降至零性能水平,真实预测模型获得正确奖励。最终实验表明,以WMSE作为损失函数可有效减少模式崩溃并提升模型性能。
原文摘要 · Abstract (English)
Recent benchmarks reveal that models for single-cell perturbation response are often outperformed by simply predicting the dataset mean. We trace this anomaly to a metric artifact: control-referenced deltas and unweighted error metrics reward mode collapse whenever the control is biased or the biological signal is sparse. Large-scale \textit{in silico} simulations and analysis of two real-world perturbation datasets confirm that shared reference shifts, not genuine biological change, drives high performance in these evaluations. We introduce differentially expressed gene (DEG)-aware metrics, weighted mean-squared error (WMSE) and weighted delta $R^{2}$ ($R^{2}_{w}(Δ)$) with respect to all perturbations, that measure error in niche signals with high sensitivity. We further introduce negative and positive performance baselines to calibrate these metrics. With these improvements, the mean baseline sinks to null performance while genuine predictors are correctly rewarded. Finally, we show that using WMSE as a loss function reduces mode collapse and improves model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。