用大模型自动选最优基因簇分组,让单细胞分析更客观。
HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets
- 用大模型生成基因集假设并打分,替代人工判断
- 通过簇内一致性和簇间区分度计算得分,优化分组粒度
- 在K562数据上比传统指标更贴近已知通路
大规模单细胞与Perturb-seq研究通常需对细胞聚类,并用基因本体(GO)术语标注每个簇以揭示潜在生物程序。但聚类分辨率选择和功能注释均依赖主观经验与专家判断。本文提出HYPOGENEAGENT,一个基于大语言模型(LLM)的框架,将聚类注释转化为可量化的优化任务。首先,作为基因集分析师的LLM分析每个基因程序或扰动模块内容,生成带置信度评分的GO假设排序列表;随后,使用句向量模型嵌入预测描述,计算成对余弦相似度,由代理评审团评估(i)簇内一致性(相同簇内部相似度高),称作簇内一致度;(ii)簇间区分度(不同簇间相似度低),称作簇间分离度。两者结合生成代理衍生的分辨率分数,在簇兼具内在一致性和互斥性时达到最大。在公开的K562 CRISPRi Perturb-seq数据集上的初步测试显示,该分辨率分数所选的聚类粒度与已知通路高度吻合,优于传统指标如轮廓系数、模块度分数。这些结果确立了大模型代理作为聚类分辨率与功能注释的客观评判者,为单细胞多组学研究提供了全自动、上下文感知的解析管道。
原文摘要 · Abstract (English)
Large-scale single-cell and Perturb-seq investigations routinely involve clustering cells and subsequently annotating each cluster with Gene-Ontology (GO) terms to elucidate the underlying biological programs. However, both stages, resolution selection and functional annotation, are inherently subjective, relying on heuristics and expert curation. We present HYPOGENEAGENT, a large language model (LLM)-driven framework, transforming cluster annotation into a quantitatively optimizable task. Initially, an LLM functioning as a gene-set analyst analyzes the content of each gene program or perturbation module and generates a ranked list of GO-based hypotheses, accompanied by calibrated confidence scores. Subsequently, we embed every predicted description with a sentence-embedding model, compute pair-wise cosine similarities, and let the agent referee panel score (i) the internal consistency of the predictions, high average similarity within the same cluster, termed intra-cluster agreement (ii) their external distinctiveness, low similarity between clusters, termed inter-cluster separation. These two quantities are combined to produce an agent-derived resolution score, which is maximized when clusters exhibit simultaneous coherence and mutual exclusivity. When applied to a public K562 CRISPRi Perturb-seq dataset as a preliminary test, our Resolution Score selects clustering granularities that exhibit alignment with known pathway compared to classical metrics such silhouette score, modularity score for gene functional enrichment summary. These findings establish LLM agents as objective adjudicators of cluster resolution and functional annotation, thereby paving the way for fully automated, context-aware interpretation pipelines in single-cell multi-omics studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。