arXiv:2507.21980cs.CL2025-07被引 2

用大模型仅凭环境元数据预测微生物分类与病原风险

Predicting Microbial Ontology and Pathogen Risk from Environmental Metadata with Large Language Models

  • 用大语言模型分析环境元数据进行微生物分类
  • 在零样本和少样本下准确预测大肠杆菌污染风险
  • 跨不同地区和数据格式表现出强泛化能力

传统机器学习模型在仅具备元数据的微生物组研究中难以泛化,尤其在小样本或标签格式异质的数据集中表现不佳。本文探索使用大语言模型(LLMs)对微生物样本进行分类,如EMPO 3等生物本体类别,并仅基于环境元数据预测病原污染风险,特别是大肠杆菌(E. Coli)的存在。我们评估了ChatGPT-4o、Claude 3.7 Sonnet、Grok-3和LLaMA 4在零样本与少样本设置下的表现,与随机森林等传统模型在多个真实数据集上进行对比。结果表明,LLMs不仅在本体分类任务中超越基线模型,还在污染风险预测上展现出强大能力,且能有效跨站点和元数据分布泛化。这些发现表明,大语言模型可有效处理稀疏、异构的生物元数据,为环境微生物学与生物监测提供一种有前景的纯元数据方法。

原文摘要 · Abstract (English)

Traditional machine learning models struggle to generalize in microbiome studies where only metadata is available, especially in small-sample settings or across studies with heterogeneous label formats. In this work, we explore the use of large language models (LLMs) to classify microbial samples into ontology categories such as EMPO 3 and related biological labels, as well as to predict pathogen contamination risk, specifically the presence of E. Coli, using environmental metadata alone. We evaluate LLMs such as ChatGPT-4o, Claude 3.7 Sonnet, Grok-3, and LLaMA 4 in zero-shot and few-shot settings, comparing their performance against traditional models like Random Forests across multiple real-world datasets. Our results show that LLMs not only outperform baselines in ontology classification, but also demonstrate strong predictive ability for contamination risk, generalizing across sites and metadata distributions. These findings suggest that LLMs can effectively reason over sparse, heterogeneous biological metadata and offer a promising metadata-only approach for environmental microbiology and biosurveillance applications.

大模型微生物元数据病原预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。