提出新框架分离基因表达中的扰动信号,提升单细胞扰动预测精度。
What Makes a Representation Good for Single-Cell Perturbation Prediction?

- 显式分离扰动特异信号与不变结构,避免信息混淆。
- 在多种测试中表现最优,尤其在组合扰动外推上提升显著。
- 适合关注可解释性与泛化能力的生物医学研究者。
单细胞扰动建模对理解细胞对基因扰动的响应至关重要。然而,现有方法(从因果表征学习到基础模型)常面临一个被忽视的挑战:基因表达主要受扰动无关信息主导,而扰动特异性信号本身极为稀疏。这导致学习到的表征要么混淆不变与扰动特异性信息,产生虚假且不可泛化的预测器;要么完全抑制扰动特异性信号,使其无法用于预测。为此,我们提出 PerturbedVAE,一种通用框架,旨在解决这一信号失衡问题。该框架显式分离扰动特异性信息与主导的不变结构,并恢复因果表征以有效利用此类信息进行预测。我们进一步提供可识别性分析,刻画了稀疏扰动效应可被可靠恢复的条件,从而明确了在这些条件下框架的具体实现方式。实证表明,PerturbedVAE 在广泛使用的基准数据集上多个评估设置中达到当前最优性能,尤其在外分布组合扰动预测中取得显著提升,并揭示出可解释的扰动-响应程序。
原文摘要 · Abstract (English)
Single-cell perturbation modeling is fundamental for understanding and predicting cellular responses to genetic perturbations. However, existing approaches, from causal representation learning to foundation models, often struggle with an overlooked challenge: gene expression is dominated by perturbation-invariant information, while perturbation-specific signals are intrinsically sparse. As a result, learned representations either entangle invariant and perturbation-specific information, leading to spurious and non-generalizable predictors, or suppress perturbation-specific signals altogether, rendering them ineffective for prediction. To address this, we propose PerturbedVAE, a general framework designed to resolve this signal imbalance. The framework explicitly separates perturbation-specific information from dominant invariant structure and recovers causal representations to effectively utilize such information for prediction. We further provide an identifiability analysis that characterizes the conditions under which sparse perturbation effects can be reliably recovered, thereby clarifying how the framework can be concretely specified under such conditions. Empirically, PerturbedVAE achieves state-of-the-art performance on a widely used benchmark across multiple evaluation settings, yielding significant gains on out-of-distribution combinatorial predictions and uncovering interpretable perturbation-response programs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。