揭示生物大模型内部基因表达的几何拓扑结构是否真实有效
What Topological and Geometric Structure Do Biological Foundation Models Learn? Evidence from 141 Hypotheses
- 用AI自动提出并验证141个拓扑假设,含多重空模型对照
- 12层变压器中11层以上显示显著非平凡拓扑,免疫组织信号最强
- 不同模型共享全局结构但不一致定位基因,适合生物机制研究者
当scGPT和Geneformer等生物基础模型处理单细胞基因表达时,其内部表征中的几何与拓扑结构是什么?是否具有生物学意义而非训练伪影?我们通过自主的大规模假设筛选——一个由AI驱动的执行-头脑风暴循环——在52轮迭代中提出了、测试并优化了141项关于持久同调、流形距离、跨模型对齐、社区结构和定向拓扑的假设,均设有明确的零假设控制和分离基因池划分。三个主要发现:第一,模型学习到真实的几何结构。基因嵌入邻域表现出非平凡拓扑,12个Transformer层中有11层在最弱领域仍显著(p < 0.05),另两个领域全部显著;多层级距离层次表明流形感知度量优于欧氏距离识别调控基因对,图社区划分可追踪已知转录因子靶点关系。第二,该结构在独立训练模型间共享。scGPT与Geneformer的CCA对齐达到0.80的典型相关性,基因检索准确率达72%,但19种方法均无法可靠恢复基因级对应关系。模型在基因空间全局形状上一致,但在具体基因位置上不一致。第三,结构更局部化。在所有零假设族的严格控制下,显著信号集中于免疫组织,而肺及外肺信号大幅减弱。
原文摘要 · Abstract (English)
When biological foundation models such as scGPT and Geneformer process single-cell gene expression, what geometric and topological structure forms in their internal representations? Is that structure biologically meaningful or a training artifact, and how confident should we be in such claims? We address these questions through autonomous large-scale hypothesis screening: an AI-driven executor-brainstormer loop that proposed, tested, and refined 141 geometric and topological hypotheses across 52 iterations, covering persistent homology, manifold distances, cross-model alignment, community structure, and directed topology, all with explicit null controls and disjoint gene-pool splits. Three principal findings emerge. First, the models learn genuine geometric structure. Gene embedding neighborhoods exhibit non-trivial topology, with persistent homology significant in 11 of 12 transformer layers at p < 0.05 in the weakest domain and 12 of 12 in the other two. A multi-level distance hierarchy shows that manifold-aware metrics outperform Euclidean distance for identifying regulatory gene pairs, and graph community partitions track known transcription factor target relationships. Second, this structure is shared across independently trained models. CCA alignment between scGPT and Geneformer yields canonical correlation of 0.80 and gene retrieval accuracy of 72 percent, yet none of 19 tested methods reliably recover gene-level correspondences. The models agree on the global shape of gene space but not on precise gene placement. Third, the structure is more localized than it first appears. Under stringent null controls applied across all null families, robust signal concentrates in immune tissue, while lung and external lung signals weaken substantially.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。