arXiv:2509.10575q-bio.GNcs.AI2025-09被引 1

用轻量开源模型实现基因集分析的精准推理。

Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis

  • 通过数据增强让小模型学会分步推理
  • 在1508个基因集中表现媲美商用大模型
  • 跨数据源泛化能力强,适合生物研究者

基因集分析(GSA)是揭示基因群分子功能的基础方法。近期基于大语言模型(LLM)的方法能为基因集标注生物功能并提供连贯解释,但多数研究依赖专有模型,虽性能更优却存在成本与数据隐私问题。此外,尚未有研究探索先进推理策略在GSA中的应用。为此,我们提出Gene-R1,一种数据增强的学习框架,赋予轻量级开源LLM针对GSA任务的逐步推理能力。在1,508个分布内基因集上的实验表明,Gene-R1取得显著性能提升,达到商用LLM水平;在106个分布外基因集上,其表现与商用及大规模模型相当,展现出对多样化基因来源的强泛化能力。

原文摘要 · Abstract (English)

The gene set analysis (GSA) is a foundational approach for uncovering the molecular functions associated with a group of genes. Recently, LLM-powered methods have emerged to annotate gene sets with biological functions together with coherent explanatory insights. However, existing studies primarily focus on proprietary models, which have been shown to outperform their open-source counterparts despite concerns over cost and data privacy. Furthermore, no research has investigated the application of advanced reasoning strategies to the GSA task. To address this gap, we introduce Gene-R1, a data-augmented learning framework that equips lightweight and open-source LLMs with step-by-step reasoning capabilities tailored to GSA. Experiments on 1,508 in-distribution gene sets demonstrate that Gene-R1 achieves substantial performance gains, matching commercial LLMs. On 106 out-of-distribution gene sets, Gene-R1 performs comparably to both commercial and large-scale LLMs, exhibiting robust generalizability across diverse gene sources.

基因分析轻量模型推理能力开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。