arXiv:2505.00017cs.CLcs.AI2025-05

用知识图谱增强大模型,让单细胞注释更准更自动

ReCellTy: Domain-Specific Knowledge Graph Retrieval-Augmented LLMs Reasoning Workflow for Single-Cell Annotation

  • 构建1.8万节点生物知识图谱,关联基因与细胞类型
  • 多任务推理流程提升注释准确率,人评分提高0.21
  • 适合生物信息学研究者,解决小模型性能不足问题

随着大语言模型(LLMs)的快速发展,其在细胞类型注释中的应用日益受到关注。然而,通用大模型在此特定任务中常因缺乏外部领域知识引导而受限。为实现更精准、全自动的细胞类型注释,我们构建了一个包含18,850个生物信息节点(包括细胞类型、基因标记、特征等)和48,944条边的全局连接知识图谱,供大模型检索差异基因相关实体以重建细胞。同时设计多任务推理工作流优化注释过程。相较于通用大模型,本方法在多个组织类型上人类评估得分提升最高达0.21,语义相似度提高6.1%,更贴近人工注释的认知逻辑。同时缩小了大模型与小模型在细胞类型注释中的性能差距,为生物信息学中结构化知识融合与推理提供新范式。

原文摘要 · Abstract (English)

With the rapid development of large language models (LLMs), their application to cell type annotation has drawn increasing attention. However, general-purpose LLMs often face limitations in this specific task due to the lack of guidance from external domain knowledge. To enable more accurate and fully automated cell type annotation, we develop a globally connected knowledge graph comprising 18850 biological information nodes, including cell types, gene markers, features, and other related entities, along with 48,944 edges connecting these nodes, which is used by LLMs to retrieve entities associated with differential genes for cell reconstruction. Additionally, a multi-task reasoning workflow is designed to optimise the annotation process. Compared to general-purpose LLMs, our method improves human evaluation scores by up to 0.21 and semantic similarity by 6.1% across multiple tissue types, while more closely aligning with the cognitive logic of manual annotation. Meanwhile, it narrows the performance gap between large and small LLMs in cell type annotation, offering a paradigm for structured knowledge integration and reasoning in bioinformatics.

单细胞知识图谱大模型生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。