arXiv:2506.20598cs.AIcs.SY2025-06被引 2

用大模型构建多智能体系统,加速可持续蛋白研究

Fine-Tuning and Prompt Engineering of LLMs, for the Creation of Multi-Agent AI for Addressing Sustainable Protein Production Challenges

  • 设计双智能体系统:检索文献+提取信息
  • 微调使信息匹配度提升25%,最高达0.94
  • 适合科研人员快速整合微生物蛋白知识

全球对可持续蛋白源的需求推动了高效智能工具的发展。本研究提出一个基于检索增强生成(RAG)的多智能体AI框架,聚焦微生物蛋白生产。系统包含两个基于GPT的智能体:文献搜索智能体用于检索特定微生物菌株的相关科学文献;信息提取智能体则从检索结果中抽取生物化学信息。通过微调和提示工程两种方法优化智能体性能。两者均显著提升信息提取准确率,平均余弦相似度最高提升25%,普遍达到≥0.89;其中微调表现更优,平均得分稳定在≥0.94,而提示工程具有更低的统计不确定性。研究还开发了用户界面,并初步探索了基于化学安全性的检索功能。

原文摘要 · Abstract (English)

The global demand for sustainable protein sources has accelerated the need for intelligent tools that can rapidly process and synthesise domain-specific scientific knowledge. In this study, we present a proof-of-concept multi-agent Artificial Intelligence (AI) framework designed to support sustainable protein production research, with an initial focus on microbial protein sources. Our Retrieval-Augmented Generation (RAG)-oriented system consists of two GPT-based LLM agents: (1) a literature search agent that retrieves relevant scientific literature on microbial protein production for a specified microbial strain, and (2) an information extraction agent that processes the retrieved content to extract relevant biological and chemical information. Two parallel methodologies, fine-tuning and prompt engineering, were explored for agent optimisation. Both methods demonstrated effectiveness at improving the performance of the information extraction agent in terms of transformer-based cosine similarity scores between obtained and ideal outputs. Mean cosine similarity scores were increased by up to 25%, while universally reaching mean scores of $\geq 0.89$ against ideal output text. Fine-tuning overall improved the mean scores to a greater extent (consistently of $\geq 0.94$) compared to prompt engineering, although lower statistical uncertainties were observed with the latter approach. A user interface was developed and published for enabling the use of the multi-agent AI system, alongside preliminary exploration of additional chemical safety-based search capabilities

多智能体大模型应用可持续蛋白

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。