arXiv:2409.02864cs.AIcs.IR2024-09被引 3

用大模型整合生物工具,让科研更智能高效。

Language Model Powered Digital Biology with BRAD

  • 基于大模型构建可配置的智能代理,打通本地文件、数据库与软件
  • 支持检索增强生成与自动化分析流程,实现上下文感知的半自治操作
  • 适合生物信息学研究者快速调用工具,提升实验设计效率

大语言模型正深刻改变生物学、计算机科学与日常生活。然而,如何整合各类计算工具、数据库与文献资源,仍是生物研究的挑战。大语言模型擅长处理非结构化信息、高效检索与自动化工作流执行。为此,我们提出原型系统BRAD(Bioinformatics Retrieval Augmented Digital assistant),一个集成多种生信工具的聊天机器人与智能体系统。其核心为一个由大模型驱动的Python AI Agent,可连接本地文件系统、在线数据库及用户软件。该Agent高度可配置,支持检索增强生成、跨生信数据库搜索及软件流水线执行。通过协同整合生信工具,BRAD构建了一个具备上下文感知能力的半自主系统,超越传统大模型聊天机器人的功能边界。系统提供图形化界面,操作直观便捷。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) are transforming biology, computer science, engineering, and every day life. However, integrating the wide array of computational tools, databases, and scientific literature continues to pose a challenge to biological research. LLMs are well-suited for unstructured integration, efficient information retrieval, and automating standard workflows and actions from these diverse resources. To harness these capabilities in bioinformatics, we present a prototype Bioinformatics Retrieval Augmented Digital assistant (BRAD). BRAD is a chatbot and agentic system that integrates a variety of bioinformatics tools. The Python package implements an AI \texttt{Agent} that is powered by LLMs and connects to a local file system, online databases, and a user's software. The \texttt{Agent} is highly configurable, enabling tasks such as Retrieval-Augmented Generation, searches across bioinformatics databases, and the execution of software pipelines. BRAD's coordinated integration of bioinformatics tools delivers a context-aware and semi-autonomous system that extends beyond the capabilities of conventional LLM-based chatbots. A graphical user interface (GUI) provides an intuitive interface to the system.

大模型生物信息智能代理工具整合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。