arXiv:2512.03187cs.LG2025-12

用哈希分块加速单细胞数据异常检测,提升效率与精度

Neighborhood density estimation using space-partitioning based hashing schemes

  • 基于空间分块哈希构建轻量级数据摘要
  • 在真实数据上比现有方法快2倍以上,准确率超90%
  • 适合处理大规模流式数据的实时监测场景

本文提出FiRE/FiRE.1,一种基于数据摘要的新型算法,用于快速识别大规模单细胞RNA测序数据中的稀有细胞亚群。该方法在多个真实数据集上表现出优于当前最优技术的性能。此外,论文还提出Enhash,一种利用投影哈希实现快速且资源高效的集成学习器,可有效检测流数据中的概念漂移,在不同类型的漂移下均展现出优异的时间效率和准确性。

原文摘要 · Abstract (English)

This work introduces FiRE/FiRE.1, a novel sketching-based algorithm for anomaly detection to quickly identify rare cell sub-populations in large-scale single-cell RNA sequencing data. This method demonstrated superior performance against state-of-the-art techniques. Furthermore, the thesis proposes Enhash, a fast and resource-efficient ensemble learner that uses projection hashing to detect concept drift in streaming data, proving highly competitive in time and accuracy across various drift types.

单细胞测序异常检测流数据哈希

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。