arXiv:2512.22225q-bio.QMcs.AI2025-12

用AI从文献中自动挖掘营养素生产菌株,提升研发效率。

Literature Mining System for Nutraceutical Biosynthesis: From AI Framework to Biological Insight

  • 基于大模型和优化提示词,自动识别文献中的营养素生产菌株。
  • 构建35个营养素-菌株关联数据集,发现谷氨酸棒杆菌等主导菌种。
  • 适合微生物育种、合成生物学和精准发酵领域的研究人员使用。

从科学文献中提取结构化知识仍是营养素研究的瓶颈,尤其在识别参与化合物生物合成的微生物菌株方面。本研究提出一种针对领域适配的大语言模型系统,结合先进提示工程,实现从非结构化文本中自动化识别营养素生产菌株。通过少量样本提示和定制查询设计,系统在多种配置下表现稳健,DeepSeekV3在准确性上优于LLaMA2,尤其在包含领域特异性菌株信息时更优。研究生成了一个涵盖氨基酸、膳食纤维、植物化学物和维生素的35个营养素-菌株关联的结构化验证数据集。结果揭示了单培养与共培养体系中显著的微生物多样性,谷氨酸棒状杆菌、大肠杆菌和枯草芽孢杆菌贡献突出,同时涌现出新型合成菌群。该人工智能框架不仅提升了文献挖掘的可扩展性与可解释性,还为菌株筛选、合成生物学设计及高价值营养素的精准发酵策略提供可操作洞察。

原文摘要 · Abstract (English)

The extraction of structured knowledge from scientific literature remains a major bottleneck in nutraceutical research, particularly when identifying microbial strains involved in compound biosynthesis. This study presents a domain-adapted system powered by large language models (LLMs) and guided by advanced prompt engineering techniques to automate the identification of nutraceutical-producing microbes from unstructured scientific text. By leveraging few-shot prompting and tailored query designs, the system demonstrates robust performance across multiple configurations, with DeepSeekV3 outperforming LLaMA2 in accuracy, especially when domain-specific strain information is included. A structured and validated dataset comprising 35 nutraceutical-strain associations was generated, spanning amino acids, fibers, phytochemicals, and vitamins. The results reveal significant microbial diversity across monoculture and co-culture systems, with dominant contributions from Corynebacterium glutamicum, Escherichia coli, and Bacillus subtilis, alongside emerging synthetic consortia. This AI-driven framework not only enhances the scalability and interpretability of literature mining but also provides actionable insights for microbial strain selection, synthetic biology design, and precision fermentation strategies in the production of high-value nutraceuticals.

文献挖掘AI制药合成生物学菌株筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。