arXiv:2503.20730cs.LG2025-03中稿 · ICLR被引 1

提出新评估指标与基准,优化单细胞数据批次效应校正方法

Benchmarking and optimizing organism wide single-cell RNA alignment methods

  • 引入KNI评分统一衡量批次效应与细胞类型预测精度
  • BA-scVI通过对抗训练显著降低批次效应,性能优于现有方法
  • 验证全生物体细胞类型图谱可由单一模型构建且信息不丢失

针对单细胞RNA测序数据中批次效应校正方法的评估标准不统一问题,本文构建了包含11个(scMARK)和46个(scREF)人类scRNA研究的小型与大型基准数据集,并标准化作者标注。提出K-Neighbors Intersection(KNI)评分,综合衡量批次效应消除与跨数据集细胞类型标签预测准确率。基于此,评估并优化了多种整合方法,提出批对抗单细胞变分推断(BA-scVI),通过编码器与解码器中的对抗训练抑制批次效应。结果表明,在对齐空间中细胞类型分组的粒度保持不变,支持使用单一模型构建全生物体细胞类型图谱而不丢失信息。

原文摘要 · Abstract (English)

Many methods have been proposed for removing batch effects and aligning single-cell RNA (scRNA) datasets. However, performance is typically evaluated based on multiple parameters and few datasets, creating challenges in assessing which method is best for aligning data at scale. Here, we introduce the K-Neighbors Intersection (KNI) score, a single score that both penalizes batch effects and measures accuracy at cross-dataset cell-type label prediction alongside carefully curated small (scMARK) and large (scREF) benchmarks comprising 11 and 46 human scRNA studies respectively, where we have standardized author labels. Using the KNI score, we evaluate and optimize approaches for cross-dataset single-cell RNA integration. We introduce Batch Adversarial single-cell Variational Inference (BA-scVI), as a new variant of scVI that uses adversarial training to penalize batch-effects in the encoder and decoder, and show this approach outperforms other methods. In the resulting aligned space, we find that the granularity of cell-type groupings is conserved, supporting the notion that whole-organism cell-type maps can be created by a single model without loss of information.

单细胞测序批次效应数据整合深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。