arXiv:2602.22247q-bio.GNcs.AI2026-02被引 1

揭示单细胞大模型如何用几何结构编码细胞生物学知识

Multi-Dimensional Spectral Geometry of Biological Knowledge in Single-Cell Transformer Representations

  • 通过63轮自动化假设检验,发现模型将基因组织成有生物学意义的坐标系
  • 主谱轴区分分泌蛋白与胞质蛋白,中间层反映分泌通路顺序
  • 关键基因调控关系、细胞类型标记物等均在低维空间清晰可辨

单细胞基础模型如scGPT学习高维基因表征,但其编码的生物学知识仍不明确。我们通过63轮自动化假设筛选(共测试183个假设),系统解码scGPT内部表征的几何结构,发现模型将基因组织成有结构的生物坐标系,而非模糊特征空间。主导谱轴按亚细胞定位分离基因,分泌蛋白位于一端,胞质蛋白位于另一端。中间变换层短暂编码线粒体和内质网区室,顺序呼应细胞分泌通路。正交轴编码蛋白质互作网络,对实验测得互作强度的拟合度极高(Spearman rho = 1.000,n = 5个STRING置信度五分位,p = 0.017)。在六维紧凑谱子空间中,模型可区分转录因子与其靶基因(AUROC = 0.744,所有12层均显著)。早期层保留具体调控关系,深层则压缩为调节者与被调节者粗略区分。抑制边在几何上更显著,B细胞主调控因子BATF和BACH2随深度向B细胞身份锚点PAX5收敛。细胞类型标记基因聚类精度高(AUROC = 0.851)。残差流几何结构与注意力模式互补。结果表明,生物变压器学习到可解释的细胞组织内部模型,对调控网络推断、药物靶点优先排序及模型审计具有意义。

原文摘要 · Abstract (English)

Single-cell foundation models such as scGPT learn high-dimensional gene representations, but what biological knowledge these representations encode remains unclear. We systematically decode the geometric structure of scGPT internal representations through 63 iterations of automated hypothesis screening (183 hypotheses tested), revealing that the model organizes genes into a structured biological coordinate system rather than an opaque feature space. The dominant spectral axis separates genes by subcellular localization, with secreted proteins at one pole and cytosolic proteins at the other. Intermediate transformer layers transiently encode mitochondrial and ER compartments in a sequence that mirrors the cellular secretory pathway. Orthogonal axes encode protein-protein interaction networks with graded fidelity to experimentally measured interaction strength (Spearman rho = 1.000 across n = 5 STRING confidence quintiles, p = 0.017). In a compact six-dimensional spectral subspace, the model distinguishes transcription factors from their target genes (AUROC = 0.744, all 12 layers significant). Early layers preserve which specific genes regulate which targets, while deeper layers compress this into a coarser regulator versus regulated distinction. Repression edges are geometrically more prominent than activation edges, and B-cell master regulators BATF and BACH2 show convergence toward the B-cell identity anchor PAX5 across transformer depth. Cell-type marker genes cluster with high fidelity (AUROC = 0.851). Residual-stream geometry encodes biological structure complementary to attention patterns. These results indicate that biological transformers learn an interpretable internal model of cellular organization, with implications for regulatory network inference, drug target prioritization, and model auditing.

单细胞可解释性图神经网络基因调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。