用AI自动标注单细胞测序数据,准确率达82.5%。
DeepSeq: High-Throughput Single-Cell RNA Sequencing Data Labeling via Web Search-Augmented Agentic Generative AI Foundation Models
- 用具备实时网络搜索能力的智能体模型自动标注数据
- 标注准确率最高达82.5%,大幅提升处理效率
- 适合大规模生物数据标注与临床诊断研究者
生成式AI基础模型为处理结构化生物数据带来变革性潜力,尤其在单细胞RNA测序领域,数据集正迅速扩展至数十亿细胞规模。我们提出利用具备实时网络搜索功能的智能体基础模型,自动化实验数据标注,最高实现82.5%的准确率。该方法有效缓解了结构化组学数据监督学习中的标注瓶颈,无需人工校对即可显著提升标注吞吐量并减少人为误差。本方法支持构建虚拟细胞基础模型,可应用于细胞类型识别与扰动预测等下游任务。随着数据量增长,这些模型有望超越人类标注性能,为大规模扰动筛选提供可靠推断支持。该应用体现了健康监测与诊断领域的特定创新,契合人类细胞图谱(Human Cell Atlas)和人类肿瘤图谱网络(Human Tumor Atlas Network)等项目目标。
原文摘要 · Abstract (English)
Generative AI foundation models offer transformative potential for processing structured biological data, particularly in single-cell RNA sequencing, where datasets are rapidly scaling toward billions of cells. We propose the use of agentic foundation models with real-time web search to automate the labeling of experimental data, achieving up to 82.5% accuracy. This addresses a key bottleneck in supervised learning for structured omics data by increasing annotation throughput without manual curation and human error. Our approach enables the development of virtual cell foundation models capable of downstream tasks such as cell-typing and perturbation prediction. As data volume grows, these models may surpass human performance in labeling, paving the way for reliable inference in large-scale perturbation screens. This application demonstrates domain-specific innovation in health monitoring and diagnostics, aligned with efforts like the Human Cell Atlas and Human Tumor Atlas Network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。