arXiv:2603.10261cs.LGq-bio.CB2026-03

从scGPT中挖出高效造血算法,无需重训练即可精准排序细胞发育路径。

Discovery of a Hematopoietic Manifold in scGPT Yields a Method for Extracting Performant Algorithms from Biological Foundation Model Internals

  • 通过三阶段方法从注意力权重提取可独立运行的造血算法
  • 在88个供体分割测试中性能超越多种主流模型,关键亚型分类准确率超0.95
  • 速度快34.5倍、参数少1000倍,适合资源受限场景快速部署

我们首次发现并提取了来自单细胞基础模型scGPT的紧凑造血算法,据知是首个通过机制可解释性从基础模型中提取出的生物学有用且具有竞争力的算法。scGPT内部编码了具显著发育分支结构的造血流形,经严格非重叠的Tabula Sapiens外部数据集验证,并通过冻结头部零样本迁移至独立多供体免疫数据集确认。为分离该几何结构,我们提出一种通用三阶段提取方法:直接从冻结注意力权重导出算子、引入轻量级学习适配器、再设计任务专用读出层,最终生成无需目标数据集微调的独立算法。在88-供体留出基准测试中,该算法在伪时间深度排序上表现最强,关键亚型终点(CD4/CD8 AUROC 0.867,单核/巨噬细胞 AUROC 0.951)领先于scVI、Palantir、DPT、CellTypist、PCA及原始表达基线。相比用三层MLP探测冻结scGPT嵌入,提取的头部在6/8分类任务上显著更优,且完成全部12-分割评估快34.5倍,仅需约1000倍更少可训练参数。导出算子从三个池化注意力头压缩为单头无显著损失,进一步降至秩64代理。对紧凑算子的机制可解释性分析揭示一个四因子核心,解释66.2%消融影响,分别对应T/淋巴系、B/浆细胞、粒细胞及单核/巨噬细胞基因程序。补充的第二流形验证(细胞间通信几何)表明该方法泛化能力不限于造血过程。

原文摘要 · Abstract (English)

We report the discovery and extraction of a compact hematopoietic algorithm from the single-cell foundation model scGPT, to our knowledge the first biologically useful, competitive algorithm extracted from a foundation model via mechanistic interpretability. We show that scGPT internally encodes a compact hematopoietic manifold with significant developmental branch structure, validated on a strict non-overlap Tabula Sapiens external panel and confirmed via frozen-head zero-shot transfer to an independent multi-donor immune panel. To isolate this geometry, we introduce a general three-stage extraction method consisting of direct operator export from frozen attention weights, a lightweight learned adaptor, and a task-specific readout, producing a standalone algorithm without target-dataset retraining. In 88-split donor-holdout benchmarks against scVI, Palantir, DPT, CellTypist, PCA, and raw-expression baselines, the extracted algorithm achieves the strongest pseudotime-depth ordering and leads on key subtype endpoints (CD4/CD8 AUROC 0.867, mono/macro AUROC 0.951). Compared to standard probing of frozen scGPT embeddings with a 3-layer MLP, the extracted head is BH-significantly better on 6/8 classification endpoints while completing a full 12-split evaluation campaign 34.5x faster with approximately 1000x fewer trainable parameters. The exported operator compresses from three pooled attention heads to a single head without statistically significant loss, and further to a rank-64 surrogate. Mechanistic interpretability of the compact operator reveals a concentrated four-factor core explaining 66.2% of ablation impact, with factors resolving into explicit T/lymphoid, B/plasma, granulocytic, and monocyte/macrophage gene programs. A supplementary second-manifold validation (intercellular communication geometry) confirms that the extraction method generalizes beyond hematopoiesis.

单细胞算法提取可解释性造血建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。