arXiv:2410.11468cs.LG2024-10被引 7

用稀疏自编码器解析单细胞基因表达模型中的生物信号

Can sparse autoencoders make sense of gene expression latent variable models?

  • 将稀疏自编码器用于高维生物数据,实现可解释的潜在特征分解
  • 在模拟数据中验证了其还原真实生成变量的能力,且能发现细微生物信号
  • 提出scFeatureLens工具,自动关联基因集与特征,助力大规模生物学研究

稀疏自编码器(SAEs)近期被用于大型语言模型中提取可解释的潜在特征。通过将密集嵌入投影到更高维且稀疏的空间,学习到的特征更解耦、更易解释。本文探索了SAEs在复杂高维生物数据嵌入分解中的潜力。利用模拟数据,系统分析了SAEs在提取潜在空间真实生成变量方面的有效性、超参数影响及局限性。应用于预训练单细胞模型的嵌入结果表明,SAEs能识别并调控关键生物过程,甚至揭示原本可能被忽略的细微生物信号。此外,本文提出scFeatureLens,一种自动化可解释性方法,通过将SAE特征与基因集对应的生物概念关联,支持单细胞基因表达模型的大规模分析与假说生成。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have lately been used to uncover interpretable latent features in large language models. By projecting dense embeddings into a much higher-dimensional and sparse space, learned features become disentangled and easier to interpret. This work explores the potential of SAEs for decomposing embeddings in complex and high-dimensional biological data. Using simulated data, it outlines the efficacy, hyperparameter landscape, and limitations of SAEs when it comes to extracting ground truth generative variables from latent space. The application to embeddings from pretrained single-cell models shows that SAEs can find and steer key biological processes and even uncover subtle biological signals that might otherwise be missed. This work further introduces scFeatureLens, an automated interpretability approach for linking SAE features and biological concepts from gene sets to enable large-scale analysis and hypothesis generation in single-cell gene expression models.

单细胞测序稀疏自编码器可解释性生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。