用语言模型思路解析细胞数据,揭示高维空间中的结构规律
The cell as a token: high-dimensional geometry in language models and cell embeddings
- 将细胞视为'令牌',在高维向量空间中建模其活性特征
- 发现低维流形结构增强嵌入空间的稳定性和可解释性
- 借鉴大模型可解释性技术,助力构建虚拟细胞图谱
单细胞测序技术将细胞映射到编码其内部活动的高维空间。最近提出的虚拟细胞模型基于大规模细胞图谱预训练学习到的模式,丰富了细胞表征。本文探讨自然语言嵌入结构的进展如何启发对单细胞数据集的分析。两领域均通过将非结构化数据划分为嵌入在高维向量空间中的'令牌'进行处理。我们讨论令牌上下文如何影响嵌入空间的几何结构,以及低维流形如何塑造该空间的鲁棒性与可解释性。文章强调语言基础模型的新进展(如可解释性探针和上下文推理)可为构建细胞图谱和训练虚拟细胞模型提供启示。
原文摘要 · Abstract (English)
Single-cell sequencing technology maps cells to a high-dimensional space encoding their internal activity. Recently-proposed virtual cell models extend this concept, enriching cells' representations based on patterns learned from pretraining on vast cell atlases. This review explores how advances in understanding the structure of natural language embeddings informs ongoing efforts to analyze single-cell datasets. Both fields process unstructured data by partitioning datasets into tokens embedded within a high-dimensional vector space. We discuss how the context of tokens influences the geometry of embedding space, and how low-dimensional manifolds shape this space's robustness and interpretation. We highlight how new developments in foundation models for language, such as interpretability probes and in-context reasoning, can inform efforts to construct cell atlases and train virtual cell models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。