首次发现单细胞转录组数据中存在类似语言模型的缩放定律。
Scaling Laws for Masked-Reconstruction Transformers on Single-Cell Transcriptomics
- 在海量单细胞数据上训练掩码重建变压器,验证缩放定律存在
- 数据充足时损失呈幂律下降,最低不可消除误差约1.44
- 数据少时模型越大效果越差,说明数据量是关键瓶颈
神经缩放定律——损失、模型规模与数据量之间的幂律关系——已在语言和视觉变压器中广泛记录,但在单细胞基因组学中仍基本未被探索。本文首次系统研究了在单细胞RNA测序(scRNA-seq)数据上训练的掩码重建变压器的缩放行为。基于CELLxGENE Census中的表达谱,构建了两种实验场景:数据丰富场景(512个高变基因,20万细胞)和数据受限场景(1,024个基因,1万细胞)。在七种模型规模(参数量从533到3.4×10⁸)下,对验证集均方误差(MSE)拟合参数化缩放定律。数据丰富场景显示出明显的幂律缩放,不可消除损失底限约为1.44;而数据受限场景几乎无缩放效应,表明当数据稀缺时模型容量并非主要约束。结果表明,当数据充足时,单细胞转录组学中确实存在类自然语言处理的缩放定律,并确认数据-参数比是决定缩放行为的关键因素。初步将数据丰富场景的渐近底限转换为信息论单位,估算出每个掩码基因位置约含2.30比特熵。文章讨论了对单细胞基础模型设计的启示,并指出需进一步测量以精炼该熵估计。
原文摘要 · Abstract (English)
Neural scaling laws -- power-law relationships between loss, model size, and data -- have been extensively documented for language and vision transformers, yet their existence in single-cell genomics remains largely unexplored. We present the first systematic study of scaling behaviour for masked-reconstruction transformers trained on single-cell RNA sequencing (scRNA-seq) data. Using expression profiles from the CELLxGENE Census, we construct two experimental regimes: a data-rich regime (512 highly variable genes, 200,000 cells) and a data-limited regime (1,024 genes, 10,000 cells). Across seven model sizes spanning three orders of magnitude in parameter count (533 to 3.4 x 10^8 parameters), we fit the parametric scaling law to validation mean squared error (MSE). The data-rich regime exhibits clear power-law scaling with an irreducible loss floor of c ~ 1.44, while the data-limited regime shows negligible scaling, indicating that model capacity is not the binding constraint when data are scarce. These results establish that scaling laws analogous to those observed in natural language processing do emerge in single-cell transcriptomics when sufficient data are available, and they identify the data-to-parameter ratio as a critical determinant of scaling behaviour. A preliminary conversion of the data-rich asymptotic floor to information-theoretic units yields an estimate of approximately 2.30 bits of entropy per masked gene position. We discuss implications for the design of single-cell foundation models and outline the additional measurements needed to refine this entropy estimate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。