arXiv:2603.01752cs.LGq-bio.CB2026-03被引 1

揭示单细胞大模型的因果计算结构,发现抑制主导与跨模型一致性。

Causal Circuit Tracing Reveals Distinct Computational Architectures in Single-Cell Foundation Models: Inhibitory Dominance, Biological Coherence, and Cross-Model Convergence

  • 通过消融SAE特征并测量下游响应,实现因果电路追踪。
  • 模型具53%生物一致性、65%-89%抑制主导性,且跨架构稳定。
  • 发现疾病相关模块更易成为跨模型共识,适合生物机制研究者。

稀疏自编码器(SAEs)可将大模型激活分解为可解释特征,但生物大模型中特征间的因果交互仍未知。本文提出因果电路追踪方法,通过消融SAE特征并测量下游响应,应用于Geneformer V2-316M和scGPT全人类模型,在四种条件下分析了96,892条边和80,191次前向传播。两个模型均表现出约53%的生物一致性与65%至89%的抑制主导性,且不受架构和细胞类型影响。scGPT产生更强效应(平均绝对d=1.40,高于Geneformer的1.05),动态更均衡。跨模型共识识别出1,142个保守域对(富集10.6倍,p<0.001),疾病相关域的共识概率高3.59倍。基因层面CRISPRi验证显示56.4%的方向准确性,证实共表达而非因果编码。

原文摘要 · Abstract (English)

Motivation: Sparse autoencoders (SAEs) decompose foundation model activations into interpretable features, but causal feature-to-feature interactions across network depth remain unknown for biological foundation models. Results: We introduce causal circuit tracing by ablating SAE features and measuring downstream responses, and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96,892 edges, 80,191 forward passes). Both models show approximately 53 percent biological coherence and 65 to 89 percent inhibitory dominance, invariant to architecture and cell type. scGPT produces stronger effects (mean absolute d = 1.40 vs. 1.05) with more balanced dynamics. Cross-model consensus yields 1,142 conserved domain pairs (10.6x enrichment, p < 0.001). Disease-associated domains are 3.59x more likely to be consensus. Gene-level CRISPRi validation shows 56.4 percent directional accuracy, confirming co-expression rather than causal encoding.

单细胞模型因果推理生物一致性神经网络可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。