arXiv:2608.00152cs.LGcs.AI2026-08被引 1

CRISPR基因编辑预测中,锁定评估流程揭示了模型在分布外失效的真相。

Can We Trust In-Distribution Success? Locked Evaluation Reveals Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation Prediction

  • 采用预注册冻结协议,确保评估过程透明可复现
  • 模型在外部数据上零样本迁移失败,相关系数为-0.139和-0.267
  • 结果表明预测信号可能被样本量混淆,需区分主次证据

当领域内基准表现无法应对分布偏移,或基准结果与设计因素纠缠时,AI评估可能引致错误推断。本研究以CRISPRi扰动效应预测为例,采用冻结的Geneformer表征,在预注册、锁定的评估协议下进行测试:模型头与选择在测试前固定;外部结果标签在最终揭盲前始终隐藏;分析决策提前确定。在虚拟细胞挑战(VCC)的域内评估中,该表征显著优于维度匹配的随机特征控制(Delta R² = +0.1645,95% CI [+0.1375, +0.1920]),满足事前注册的有用性门槛。然而在两个外部筛选数据集上,其零样本迁移表现均失败(斯皮尔曼等级相关系数rho = -0.139 和 -0.267),低于随机控制。加入预定义幅度块虽提升外部表现(Delta rho = +0.032 和 +0.143),但在冻结主头条件下仍无法恢复迁移能力,两者仍为负值。一个预注册的、计数调整的最高响应二级指标在两个数据集上与结果呈正相关,但仅报告为相关性而非重建的幅度信号。最后,发现VCC终点强依赖样本量:仅用计数的线性模型已达R² = +0.4325,而四个幅度标量仅贡献R² = +0.2589;将幅度标量加入计数后,仅提升R² = +0.0017,说明多数聚合幅度信号与样本量高度重叠。本案例表明,通过锁定评估、统一测量终点、分离主次证据,可改变对AI基准推断的支持方向。

原文摘要 · Abstract (English)

AI evaluation can support the wrong inference when an in-domain benchmark success does not survive distribution shift, or when the benchmark endpoint is entangled with a design factor. We study this problem in CRISPRi perturbation-effect prediction, evaluating a frozen Geneformer representation under a locked, pre-registered protocol: heads and model selection were frozen before test evaluation; the protocol required external outcome labels to remain withheld until final unblinding; and analysis-governing decisions were fixed before the evaluations they govern. In-distribution on the Virtual Cell Challenge (VCC), the frozen representation carries measurable predictive information beyond a dimension-matched random-feature control (Delta R^2 = +0.1645, 95% CI [+0.1375, +0.1920]), satisfying the pre-registered informativeness gate required before interpreting transfer. It then fails zero-shot transfer on both external screens (Spearman rho = -0.139 and -0.267), lying below that control on each. Adding a predefined magnitude block improves the representation externally (Delta rho = +0.032 and +0.143) but, under the frozen primary head, does not rescue transfer: both remain negative. A pre-registered, count-adjusted max-response secondary is positively associated with the outcome on both screens; we report it as correlational and secondary, not as a recovered magnitude signal. Finally, the VCC endpoint is strongly sample-size associated: a count-only linear model reaches R^2 = +0.4325, versus +0.2589 for the four magnitude scalars; adding those scalars to cell count improves R^2 by only +0.0017, so much of the aggregate-magnitude signal overlaps with cell count. This case study shows how locking the evaluation, harmonizing the measured endpoint, and separating primary from secondary evidence can change the inference supported by an AI benchmark.

基因预测评估可靠性分布外泛化因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。