arXiv:2607.25288cs.LGq-bio.GN2026-07中稿 · ISRSD 2026

对比九种方法,找出深度学习何时真能提升单细胞聚类效果

When Does Deep Representation Learning Help Single-Cell Clustering? A Sensitivity-Aware Diagnostic Benchmark for Biomedical AI Pipelines

  • 构建诊断性基准,系统测试深度与经典方法在真实数据上的表现
  • 小数据用变分自编码器,中等数据用深度自编码器,大线性数据仍可用PCA
  • 揭示学习率和隐空间维度是影响结果最关键的两个参数

单细胞RNA测序(scRNA-seq)是精准医疗的核心技术,支撑联合国可持续发展目标3。无监督聚类将原始表达矩阵转化为可解释的细胞群。研究者常面临抉择:是否值得投入计算成本使用深度表示?本文在10个真实数据集(90–5,685细胞,19,046–41,480基因,4–11种细胞类型)上评估了9种聚类流程,另在7个数据集上补充了部分scVI V2对比。方法融合Optuna超参搜索、重复运行鲁棒性、Friedman/Wilcoxon-Holm/TOST检验及Sobol总阶敏感性分析。对比自编码器取得最高平均调整兰德指数(0.7872),但经Holm校正后未显著优于最强基线。每数据集分析发现三类可复现模式:变分自编码器在最小数据集有优势,深度自编码器在中等规模且含多批次或多种类结构时更优,而线性投影已捕捉主要变异时,传统PCA依然有效。Sobol指数显示学习率(ST=0.70)和隐空间维度(ST=0.56)是主导方差来源,提示调参资源应优先分配于此。贡献在于构建一个数据感知、算力敏感的生物医学AI决策框架,而非宣称某方法普适最优。

原文摘要 · Abstract (English)

Single-cell ribonucleic acid sequencing (scRNA-seq) is a foundational technology for precision-medicine workflows that contribute to United Nations Sustainable Development Goal 3 on Good Health and Well-being, and unsupervised clustering is the analytical step that turns raw expression matrices into interpretable cell populations. Practitioners therefore face a recurring engineering decision: is an additional deep representation stage worth its compute and tuning cost, or do classical principal component analysis (PCA) pipelines already suffice? We address this question with a diagnostic benchmark of nine clustering pipelines on ten real datasets (90-5,685 cells, 19,046-41,480 genes, 4-11 cell types), augmented by a partial scVI V2 specialized comparison on seven datasets. The protocol integrates Optuna hyperparameter search, repeated-run robustness, Friedman/Wilcoxon-Holm/TOST testing, and Sobol total-order sensitivity analysis. The contrastive autoencoder achieved the highest mean Adjusted Rand Index (0.7872), but Holm-corrected tests did not establish dominance over the strongest baselines. Per-dataset analysis reveals three reproducible regimes: probabilistic variational autoencoder (VAE) variants help on the smallest datasets, deep autoencoders win on mid-scale data with multi-batch or many-type structure, and classical PCA pipelines remain competitive when linear projection already captures the dominant variation. Sobol indices identify learning rate ($S_T=0.70$) and latent dimensionality ($S_T=0.56$) as the dominant variance contributors, indicating where limited tuning budgets should be allocated. The contribution is therefore a dataset-aware and compute-conscious decision framework for biomedical AI pipelines supporting sustainable healthcare analytics, rather than a universal superiority claim.

单细胞聚类深度学习敏感性分析生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。