arXiv:2606.27752cs.LG2026-06

用验证器引导强化学习,让单细胞扰动预测更符合生物真实。

PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction

论文配图:PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction
图 1 · 摘自论文原文
  • 用细胞级验证器作为奖励,优化生成模型的生物学一致性。
  • 在多个基因与化学扰动数据集上,显著提升预测准确率。
  • 适合关注单细胞预测可靠性与机制可解释性的研究者。

单细胞扰动模型可通过预测细胞对干预的转录反应,减少昂贵的湿实验筛选。尽管近期生成模型提升了群体层面的预测能力,但生成的单个细胞并未显式检查其生物学一致性。我们提出PerturbCellRL,一种基于强化学习的后训练框架,利用一组细胞级验证器作为奖励来微调预训练的单细胞转录组生成器。这些验证器定义了四项奖励:皮尔逊相关性前k名相似度、均方根误差前k名邻近度、差异表达的斯皮尔曼相关性,以及通路活性。其中通路活性验证器奖励那些通路响应与已知扰动生物学一致的细胞。我们在多个基因和化学扰动基准上评估PerturbCellRL。结果表明,在各类对齐奖励的指标及保留测试集指标上,PerturbCellRL均优于预训练的流匹配生成器,同时在群体层面指标上仍保持与顶尖方法相当的性能。这些结果表明,可信的单细胞预测应以验证器引导的生成对齐为核心,超越单纯匹配表达分布,实现对单细胞扰动效应的生物学一致性显式检验。

原文摘要 · Abstract (English)

Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. While recent generative models improve population-level prediction, individual generated cells are not explicitly checked for biological consistency. We introduce PerturbCellRL, a reinforcement learning (RL) framework that post-trains a pretrained single-cell transcriptomic generator using a suite of cell-level verifiers as rewards. These verifiers define four rewards: Pearson top-k similarity, RMSE top-k proximity, DE Spearman, and Pathway activity. The Pathway activity verifier rewards cells whose pathway responses match known perturbation biology. We evaluate PerturbCellRL on multiple genetic and chemical perturbation benchmarks. Across these benchmarks, PerturbCellRL improves over the pretrained flow-matching generator on reward-aligned evaluation metrics and a held-out evaluation metric. Moreover, PerturbCellRL remains competitive with state-of-the-art methods on population-level metrics. Together, these results frame trustworthy single-cell prediction as verifier-guided generative alignment, moving beyond matching expression distributions toward predictions whose single-cell perturbation effects are explicitly checked for biological consistency.

单细胞生成模型强化学习生物一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。