arXiv:2509.14723cs.LG2025-09

用转换器解析单细胞模型决策过程,揭示真实生物机制。

Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models

  • 训练转换器从单细胞基础模型中提取决策电路
  • 发现的电路与真实生物通路一致,具备生物学可解释性
  • 适合关注模型可解释性的生物信息学研究者

单细胞基础模型(scFMs)通过大规模转录组数据学习基因调控网络,在细胞类型注释和扰动响应预测等任务上表现优异。然而,其决策过程远不如传统方法(如差异表达分析)可解释。近期,转换器被用于从大语言模型中提取可解释决策电路。本文在先进的scFM——cell2sentence(C2S)模型上训练转换器,利用其提取C2S模型内部的决策电路。结果表明,所发现的电路对应真实生物机制,验证了转换器在揭示复杂单细胞模型内生物学合理通路方面的潜力。

原文摘要 · Abstract (English)

Single-cell foundation models (scFMs) have demonstrated state-of-the-art performance on various tasks, such as cell-type annotation and perturbation response prediction, by learning gene regulatory networks from large-scale transcriptome data. However, a significant challenge remains: the decision-making processes of these models are less interpretable compared to traditional methods like differential gene expression analysis. Recently, transcoders have emerged as a promising approach for extracting interpretable decision circuits from large language models (LLMs). In this work, we train a transcoder on the cell2sentence (C2S) model, a state-of-the-art scFM. By leveraging the trained transcoder, we extract internal decision-making circuits from the C2S model. We demonstrate that the discovered circuits correspond to real-world biological mechanisms, confirming the potential of transcoders to uncover biologically plausible pathways within complex single-cell models.

单细胞可解释性转换器基因调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。