arXiv:2601.12805q-bio.GNcs.AI2026-01KDD被引 3

构建基因级推理基准,评估大模型从基因知识推断功能的能力。

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

  • 基于19万个人类基因数据,设计54万道基因到功能的推理题。
  • 发现主流大模型在生成完整、真实功能解释上仍有显著缺陷。
  • 针对幻觉、信息不全等痛点,提供可量化的评估视角,适合生物医学研究者使用。

大型语言模型(LLMs)在生物医学研究中展现出日益增长的潜力,尤其在依赖知识的解释任务中。然而,其从基因层面知识推理出功能理解的能力——这是增强型细胞图谱解读的核心需求——仍缺乏系统评估。为此,我们提出 SciHorizon-GENE,一个大规模基因中心的基准,源自权威生物数据库。该基准整合了超过19万个人类基因的结构化知识,包含超过54万道问题,覆盖细胞类型标注、功能解释及机制分析等多种基因到功能的推理场景。受初步观察中行为模式启发,基准从四个生物学关键维度评估模型:研究关注度敏感性、幻觉倾向、答案完整性与文献影响,明确聚焦于限制其在生物解释流程中安全应用的失效模式。我们系统评估了多种前沿通用与生物医学LLMs,揭示了基因级推理能力的显著差异,以及生成忠实、完整且基于文献的功能解释仍面临持续挑战。该基准为基因尺度上分析模型行为提供了系统基础,并为模型选择与开发提供洞见,直接服务于知识增强型生物解释任务。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. However, their ability to reliably reason from gene-level knowledge to functional understanding, a core requirement for knowledge-enhanced cell atlas interpretation, remains largely underexplored. To address this gap, we introduce SciHorizon-GENE, a large-scale gene-centric benchmark constructed from authoritative biological databases. The benchmark integrates curated knowledge for over 190K human genes and comprises more than 540K questions covering diverse gene-to-function reasoning scenarios relevant to cell type annotation, functional interpretation, and mechanism-oriented analysis. Motivated by behavioral patterns observed in preliminary examinations, SciHorizon-GENE evaluates LLMs along four biologically critical perspectives: research attention sensitivity, hallucination tendency, answer completeness, and literature influence, explicitly targeting failure modes that limit the safe adoption of LLMs in biological interpretation pipelines. We systematically evaluate a wide range of state-of-the-art general-purpose and biomedical LLMs, revealing substantial heterogeneity in gene-level reasoning capabilities and persistent challenges in generating faithful, complete, and literature-grounded functional interpretations. Our benchmark establishes a systematic foundation for analyzing LLM behavior at the gene scale and offers insights for model selection and development, with direct relevance to knowledge-enhanced biological interpretation.

基因推理大模型评测生物医学AI知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。