单细胞模型中间层比最后一层更优,需按任务选提取位置。
Intermediate Layers Encode Optimal Biological Representations in Single-Cell Foundation Models
- 按任务和细胞状态选择中间层,而非默认用最后一层。
- 轨迹推断最佳层在60%深度,扰动预测最优层可差96%。
- 静息细胞中第一层反而最优,挑战抽象层级假设。
当前单细胞基础模型评估普遍使用最后一层嵌入,认为其代表最优特征空间。我们系统评估了scFoundation(100M参数)和Tahoe-X1(1.3B参数)在轨迹推断与扰动响应预测任务中的逐层表示。结果表明,最优层具有任务依赖性(轨迹在60%深度达到峰值,比最后一层高31%)和上下文依赖性(扰动最优层在不同T细胞激活状态下变化0-96%)。值得注意的是,在静息细胞中,第一层嵌入优于所有深层表示,挑战了层级抽象的假设。研究证明,'在哪提取特征'与'学什么特征'同样重要,应根据生物任务和细胞背景系统评估各层表现,而非默认使用最终层嵌入。
原文摘要 · Abstract (English)
Current single-cell foundation model benchmarks universally extract final layer embeddings, assuming these represent optimal feature spaces. We systematically evaluate layer-wise representations from scFoundation (100M parameters) and Tahoe-X1 (1.3B parameters) across trajectory inference and perturbation response prediction. Our analysis reveals that optimal layers are task-dependent (trajectory peaks at 60% depth, 31% above final layers) and context-dependent (perturbation optima shift 0-96% across T cell activation states). Notably, first-layer embeddings outperform all deeper layers in quiescent cells, challenging assumptions about hierarchical feature abstraction. These findings demonstrate that "where" to extract features matters as much as "what" the model learns, necessitating systematic layer evaluation tailored to biological task and cellular context rather than defaulting to final-layer embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。