用基因同源关系引导大模型预测非模式生物基因功能。
Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction

- 结合同源基因关系与多跳检索,从多个数据库获取证据。
- 在150个查询中99.33%成功获取相关证据,排名准确率达99%。
- 适合生物信息学研究者和基因功能注释人员使用。
非模式生物的基因功能注释仍是计算生物学中的重大挑战,其中20%-70%的测序基因功能未被明确。传统基于同源的方法成本高且依赖高序列相似性。本研究提出Homo-RAG框架,利用大语言模型结合同源引导的多跳检索与证据感知排序,通过zebrafish与人类同源基因的关系,从ZFIN、UniProt和PubMed中混合稠密与词法检索获取证据。引入证据置信度得分(ECS),融合语义相关性、实体匹配、同源信息、来源可靠性及文献关联信号,优化证据排序。在150个查询和7,200份文档上的评估显示,当权重参数lambda=0.50时,NDCG@10提升至0.9879,MRR达0.99,99.33%的查询成功获取相关证据。此外80%的检索文档为查询独有,表明证据质量补充而非替代检索相关性。结果证明Homo-RAG是可靠、可落地的基因功能预测框架。
原文摘要 · Abstract (English)
The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced genes lacking characterized functions. Traditional homology-based methods are often costly and strongly dependent on high sequence similarity. This study presents Homo-RAG, a framework for large language model-based gene function prediction that integrates homology-guided multi-hop retrieval with evidence-aware ranking. The framework exploits biological relationships between zebrafish and human orthologs to guide evidence acquisition from ZFIN, UniProt, and PubMed through hybrid dense and lexical retrieval. An Evidence Confidence Score (ECS) integrates semantic relevance, entity matching, orthology information, source reliability, and literature association signals to refine the ranking of retrieved evidence. Extensive evaluation across 150 queries and 7,200 retrieved documents shows that evidence weighting parameter of lambda=0.50 improves NDCG@10 to 0.9879 and MRR to 0.99, while retrieving relevant evidence for 99.33% of queries. Furthermore, 80% of the retrieved documents are query-exclusive, indicating that evidence quality complements rather than replaces retrieval relevance. These findings establish Homo-RAG as a practical and robust framework for reliable, evidence-grounded gene function prediction in understudied organisms. The study addresses important limitations of conventional annotation pipelines while identifying opportunities for future improvements in evidence features and attribution mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。