用可解释模型精准模拟细胞对干预的反应,助力疾病研究与药物设计。
scCBGM: Interpretable Single-Cell Counterfactual Editing

- 基于概念瓶颈架构,通过跳接和交叉协方差正则化实现细胞级可解释编辑。
- 在真实数据上表现优于现有方法,具备强组合泛化能力,预测准确率提升显著。
- 适合生物医学研究者、单细胞数据分析人员使用,尤其关注机制可解释性场景。
理解细胞表型及其对扰动的响应对疾病生物学和治疗设计至关重要。单细胞RNA测序可在细胞分辨率下进行表征,但条件组合空间庞大,难以通过实验全面映射。我们提出单细胞概念瓶颈生成模型(scCBGM),一种可解释且精确的细胞反事实编辑框架。scCBGM通过解码器跳接和交叉协方差惩罚项,将概念瓶颈架构适配于单细胞数据,促进无维度约束的特征解耦。我们进一步将其扩展至流匹配模型,支持编码-解码与生成两种模式下的概念引导编辑。为实现严格评估,我们构建了含真值反事实的合成基准。在多个真实数据集上,scCBGM在组合泛化和反事实预测方面表现优异,合成数据的细胞级验证及真实数据的群体级基准均证实其有效性。
原文摘要 · Abstract (English)
Understanding cellular phenotypes and how they respond to perturbations is critical for disease biology and therapeutic design. Single-cell RNA sequencing enables characterization at cellular resolution, yet the combinatorial space of conditions makes exhaustive experimental mapping infeasible. We introduce single-cell Concept Bottleneck Generative Models (scCBGM), a framework for interpretable and precise counterfactual editing of individual cells. scCBGM adapts concept bottleneck architectures for single-cell data through decoder skip connections and a cross-covariance penalty that promotes disentanglement without dimensional constraints. We extend the framework to flow matching models, enabling concept-guided editing in both encoding-decoding and generation regimes. To enable rigorous evaluation, we develop a synthetic benchmark with ground-truth counterfactuals. Across multiple real datasets, scCBGM demonstrates superior performance in combinatorial generalization and counterfactual prediction, supported by cell-level validation on synthetic data and population-level benchmarks on real datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。