arXiv:2509.19566cs.AIq-bio.GN2025-09被引 2

用小模型+智能框架解决基因组问答的幻觉与成本问题

Nano Bio-Agents (NBA): Small Language Model Agents for Genomics

  • 用小型语言模型搭配智能体框架分解任务、调用工具
  • 3-10B参数模型在基因组基准上达85%-97%准确率,最优组合98%
  • 计算开销低,适合资源有限的研究者使用

我们研究了参数量小于100亿的小型语言模型(SLMs)在基因组问答中的应用,通过智能体框架解决幻觉问题和计算成本挑战。所提出的纳米生物智能体(NBA)框架整合了任务分解、工具编排和对NCBI、AlphaGenome等成熟系统的API访问。结果表明,结合该框架的SLMs在性能上可媲美甚至超越现有大模型方法;最佳模型-智能体组合在GeneTuring基准测试中达到98%准确率。值得注意的是,3-10B参数的小模型在多数情况下稳定实现85%-97%的准确率,且所需计算资源远低于传统方法。这展示了在保持高鲁棒性和准确性的同时,实现效率提升、成本降低及机器学习驱动基因组工具普及的巨大潜力。

原文摘要 · Abstract (English)

We investigate the application of Small Language Models (<10 billion parameters) for genomics question answering via agentic framework to address hallucination issues and computational cost challenges. The Nano Bio-Agent (NBA) framework we implemented incorporates task decomposition, tool orchestration, and API access into well-established systems such as NCBI and AlphaGenome. Results show that SLMs combined with such agentic framework can achieve comparable and in many cases superior performance versus existing approaches utilising larger models, with our best model-agent combination achieving 98% accuracy on the GeneTuring benchmark. Notably, small 3-10B parameter models consistently achieve 85-97% accuracy while requiring much lower computational resources than conventional approaches. This demonstrates promising potential for efficiency gains, cost savings, and democratization of ML-powered genomics tools while retaining highly robust and accurate performance.

小模型基因组智能体问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。