提出新基准,解决肠道病单细胞数据分类中的捐赠者混淆问题
Donor-Aware scRNA-seq Benchmarks for IBD Classification

- 按捐赠者划分数据,避免训练与测试混用导致的虚假性能
- 肠腔分隔的细胞组成特征在溃疡性结肠炎中达到0.978的准确率
- 适合关注单细胞数据建模严谨性的生物医学研究者参考
从单细胞RNA测序(scRNA-seq)进行捐赠者级别的疾病分类,需严格采用捐赠者感知的交叉验证:随意打乱细胞分割会混淆训练与测试捐赠者,因伪重复导致性能虚高。本文构建捐赠者感知基准,评估三种特征表示在两个独立炎症性肠病(IBD)队列中的表现:中心化对数比(CLR)变换的细胞类型组成、基于GatedStructuralCFN的依赖嵌入,以及scVI变分自编码器的潜在嵌入。队列包括SCP259溃疡性结肠炎图谱(UC vs. 健康,n=30捐赠者,51种细胞类型)和Kong 2023克罗恩病图谱(CD vs. 健康,n=71捐赠者,55–68种细胞类型分布在三个肠段)。分腔段的CLR组成在SCP259上取得AUROC 0.956 ± 0.061;使用相同特征的GatedStructuralCFN达0.978 ± 0.050。在Kong队列中,CFN在结肠区表现最佳(特征筛选后0.960 ± 0.055),优于线性CLR(0.900 ± 0.100);而回肠区则由线性模型主导(CatBoost CLR 0.967 ± 0.075 vs. CFN 0.811 ± 0.164)。跨数据集迁移(CD→UC,4个共享细胞类型)使用XGBoost CLR得AUC 0.833,反向迁移接近随机水平。CFN边稳定性分析显示,分腔组成可消除全局组成因单位和带来的虚假不稳定性(Jaccard 0.026 vs. 前20项重现率1.0)。尽管在结肠区CFN数值优于线性模型(AUROC 0.960 vs. 0.900),但在每区域≤34个捐赠者的条件下,方法间差异未达统计显著性。分腔特征构建对分类性能与结构可解释性均至关重要。
原文摘要 · Abstract (English)
Donor-level disease classification from single-cell RNA sequencing (scRNA-seq) requires strict donor-aware cross-validation: naive pipelines that split cells randomly conflate training and test donors, inflating reported performance through pseudoreplication. We present a donor-aware benchmark evaluating three feature representations across two independent IBD cohorts: centered log-ratio (CLR) transformed cell-type composition, GatedStructuralCFN dependency embeddings, and scVI variational autoencoder latent embeddings. The cohorts are the SCP259 ulcerative colitis atlas (UC vs. Healthy, n=30 donors, 51 cell types) and the Kong 2023 Crohn's disease atlas (CD vs. Healthy, n=71 donors, 55-68 cell types across three intestinal regions). Compartment-stratified CLR composition achieves AUROC 0.956 +/- 0.061 on SCP259; GatedStructuralCFN on the same features achieves 0.978 +/- 0.050. In the Kong cohort, CFN achieves its best performance in the colon region (0.960 +/- 0.055 after feature filtering), exceeding linear CLR (0.900 +/- 0.100), while terminal ileum classification is dominated by linear models (CatBoost CLR 0.967 +/- 0.075 vs. CFN 0.811 +/- 0.164). Cross-dataset transfer (CD->UC, four shared cell types) achieves AUC 0.833 with XGBoost CLR; the reverse direction performs at chance. CFN edge stability analysis shows that compartment-wise composition eliminates spurious unit-sum-induced instability present in global composition (Jaccard 0.026 vs. top-20 recurrence 1.0). CFN shows a consistent numerical advantage over linear models in the colon region of CD (AUROC 0.960 vs. 0.900), though no inter-method comparison reached statistical significance at n<=34 donors per region. Compartment-aware feature construction is critical for both classification performance and structural interpretability. Code: https://github.com/Jonathan-321/sfn-scrna-study
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。