arXiv:2608.00099q-bio.GNcs.AI2026-08

用大模型自动把基因注释术语聚类成有生物学意义的高阶功能模块。

LLMBDC: Language Model for Biological Domains Oriented Clustering of Gene Ontology

论文配图:LLMBDC: Language Model for Biological Domains Oriented Clustering of Gene Ontology
图 1 · 摘自论文原文
  • 基于大模型零样本语义推理,仅靠本体信息实现无训练聚类。
  • 在阿尔茨海默病和脆性X综合征数据上,聚类准确率提升超60个百分点。
  • 适合需要可解释、可复现生物功能解读的研究者使用。

基因本体(GO)富集分析是将大规模基因组数据转化为生物见解的基础工具,但通常产生数百个冗余术语,掩盖了核心主题。现有总结工具依赖固定相似性度量(REVIGO、GOSemSim、clusterProfiler::simplify)、基因重叠指标(Metascape)或静态层级映射(GO-slim),无法融入生物背景。手动归类虽具上下文感知,但主观且耗时。亟需一种可扩展、上下文感知的框架,将GO术语聚类为可解释的高阶生物域。本文提出LLMBDC(面向生物领域的大语言模型聚类框架),无需训练,仅在推理时利用大模型的零样本语义推理与置信度评分,结合本体信息将GO术语聚类为BioDomains。在阿尔茨海默病(AD)和脆性X综合征(FXS)数据上对比六种基线方法(包括SapBERT),LLMBDC显著提升精确率、召回率及聚类性能。相较人工标注,其调整后的互信息(ARI)从9.7%提升至73.3%(AD),15.7%提升至66.6%(FXS),相应标准化互信息(NMI)从59.9%增至73.4%(AD),66.0%增至79.5%(FXS)。柯西组合检验进一步证实,聚合后的BioDomains仍保留显著的功能信号。LLMBDC为实现可扩展、可复现、可解释的上下文感知系统级解读提供了新路径,同时保持生物特异性。

原文摘要 · Abstract (English)

Gene Ontology (GO) enrichment analysis is a foundational tool for translating large-scale genomic data into biological insights, but typically yields hundreds of redundant terms that obscure overarching themes. Existing summarization tools rely on fixed similarity metrics (REVIGO, GOSemSim, clusterProfiler::simplify()), gene-overlap measures (Metascape), or static hierarchy mappings (GO-slim), and therefore cannot incorporate biological context. Manual curation provides context-aware grouping but is subjective and labor-intensive. A scalable, context-aware framework is needed to cluster GO terms into interpretable higher-order biological domains. Here we present LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology), a training-free framework that leverages zero-shot semantic reasoning of LLMs with confidence scoring to cluster GO terms into BioDomains using only ontology information at inference time. Benchmarked across Alzheimer's disease (AD) and Fragile X syndrome (FXS) against six baseline methods including SapBERT, LLMBDC achieved substantially higher precision, recall, and clustering performance. Against ground-truth annotations, LLMBDC improved ARI from 9.7% to 73.3% (AD) and from 15.7% to 66.6% (FXS) over REVIGO, with corresponding NMI gains from 59.9% to 73.4% (AD) and 66.0% to 79.5% (FXS). A Cauchy combination test further confirmed that aggregated BioDomains retained statistically significant functional signals. LLMBDC provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.

基因注释大模型聚类生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。