用小模型+多智能体让生物信息分析更简单,本地运行还能用私有数据。
BioAgents: Democratizing Bioinformatics Analysis with Multi-Agent Systems
- 基于微调的小模型和检索增强生成,构建可本地部署的多智能体系统。
- 在基因组概念任务上表现接近人类专家,代码生成能力待提升。
- 适合想自主分析数据但缺乏编程或领域知识的研究人员。
构建端到端生物信息学工作流需要跨领域的专业知识,这对初学者和资深研究者都构成挑战,因其要求同时掌握基因组学概念与计算技术。尽管大语言模型提供一定帮助,但在执行复杂生物信息学任务时常缺乏细致指导,且需昂贵算力以达高性能。为此,我们提出基于小语言模型、在生物信息学数据上微调并结合检索增强生成(RAG)的多智能体系统——BioAgents。该系统支持本地运行与个性化,可使用私有数据。实验显示其在概念性基因组任务上的表现与人类专家相当,并建议未来增强代码生成能力。
原文摘要 · Abstract (English)
Creating end-to-end bioinformatics workflows requires diverse domain expertise, which poses challenges for both junior and senior researchers as it demands a deep understanding of both genomics concepts and computational techniques. While large language models (LLMs) provide some assistance, they often fall short in providing the nuanced guidance needed to execute complex bioinformatics tasks, and require expensive computing resources to achieve high performance. We thus propose a multi-agent system built on small language models, fine-tuned on bioinformatics data, and enhanced with retrieval augmented generation (RAG). Our system, BioAgents, enables local operation and personalization using proprietary data. We observe performance comparable to human experts on conceptual genomics tasks, and suggest next steps to enhance code generation capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。