ELISA让单细胞数据生成可解释的生物学假设,打通基因表达与自然语言理解的桥梁。
ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics
- 融合基因表达嵌入与语义检索,实现多模态查询响应
- 在6个数据集上显著优于基线方法,基因签名查询提升5.98倍效应量
- 适合生物学家探索单细胞数据,生成可验证的科学假说
将单细胞RNA测序(scRNA-seq)数据转化为机制性生物学假设仍是关键瓶颈,因智能体系统缺乏对转录组表示的直接访问,而表达基础模型对自然语言不透明。本文提出ELISA(Embedding-Linked Interactive Single-cell Agent),一个可解释框架,整合scGPT表达嵌入、基于BioBERT的语义检索与大语言模型(LLM)推理,支持交互式单细胞发现。自动查询分类器根据输入类型(基因特征、自然语言概念或混合)路由至不同处理流程。集成分析模块在嵌入数据上直接完成:60+通路活性评分、280+配体-受体互作预测、条件感知比较分析及细胞类型比例估计。在炎症性肺病、儿童与成人癌症、类器官、健康组织和神经发育等6个异质数据集上,ELISA显著优于CellWhisperer、BM25及随机基线(联合置换检验,$p < 2 imes10^{-5}$),尤其在基因签名查询中表现突出(MRR Cohen's $d = 5.98$)。ELISA复现已有文献发现(平均复合得分0.88),并通过基于证据的LLM推理生成候选假说,弥合转录组探索与生物学发现之间的鸿沟。
原文摘要 · Abstract (English)
Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language. Here we introduce ELISA (Embedding-Linked Interactive Single-cell Agent), an interpretable framework that unifies scGPT expression embeddings with BioBERT-based semantic retrieval and LLM-mediated interpretation for interactive single-cell discovery. An automatic query classifier routes inputs to gene marker scoring, semantic matching, or reciprocal rank fusion pipelines depending on whether the query is a gene signature, natural language concept, or mixture of both. Integrated analytical modules perform pathway activity scoringacross 60+ gene sets, ligand-receptor interaction prediction using 280+ curated pairs, condition-aware comparative analysis, and cell-type proportion estimation all operating directly on embedded data without access to the original count matrix. Benchmarked across six diverse scRNA-seq datasets spanning inflammatory lung disease, pediatric and adult cancers, organoid models, healthy tissue, and neurodevelopment, ELISA significantly outperforms CellWhisperer, a classical lexical retriever (BM25), and a random baseline in cell type retrieval (combined permutation test, $p < 2\times10^{-5}$ for each), with particularly large gains on gene-signature queries (Cohen's $d = 5.98$ for MRR). ELISA replicates published biological findings (mean composite score 0.88), and generates candidate hypotheses through grounded LLM reasoning, bridging the gap between transcriptomic data exploration and biological discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。