arXiv:2604.05775cs.CLq-bio.GN2026-04ACL被引 1

测试大模型读不懂病毒基因组,但能猜出宿主和类型。

PhageBench: Can LLMs Understand Raw Bacteriophage Genomes?

论文配图:PhageBench: Can LLMs Understand Raw Bacteriophage Genomes?
图 1 · 摘自论文原文
  • 用专家工作流程设计数据集,评测大模型理解病毒基因序列能力。
  • 模型在识别病毒片段和预测宿主上表现优于随机猜测。
  • 对长距离依赖和精细功能定位任务仍不擅长,适合生物信息研究者关注。

噬菌体被称为生物圈的‘暗物质’,在调控微生物生态和替代抗生素方面具有关键作用。准确解析其基因组具有重要科学与应用价值。尽管通用大语言模型在理解生物文本方面表现出色,但它们直接解析原始核苷酸序列并进行生物学推理的能力仍待探索。为此,我们提出PhageBench,首个模仿生物信息学专家工作流程的基准,用于评估大模型对噬菌体基因组的理解能力。数据集包含5,600个高质量样本,涵盖筛查、质控和表型注释三个阶段的五项核心任务。对八种大模型的评估显示,通用推理模型在噬菌体连续序列识别和宿主预测任务中显著优于随机基线,展现出基因组理解的潜力。然而,在涉及长程依赖和精细功能定位的复杂推理任务中仍存在明显局限。这些发现凸显了开发具备更强推理能力的下一代生物序列模型的必要性。

原文摘要 · Abstract (English)

Bacteriophages, often referred to as the dark matter of the biosphere, play a critical role in regulating microbial ecosystems and in antibiotic alternatives. Thus, accurate interpretation of their genomes holds significant scientific and practical value. While general-purpose Large Language Models (LLMs) excel at understanding biological texts, their ability to directly interpret raw nucleotide sequences and perform biological reasoning remains underexplored. To address this, we introduce PhageBench, the first benchmark designed to evaluate phage genome understanding by mirroring the workflow of bioinformatics experts. The dataset contains 5,600 high-quality samples covering five core tasks across three stages: Screening, Quality Control, and Phenotype Annotation. Our evaluation of eight LLMs reveals that general-purpose reasoning models significantly outperform random baselines in phage contig identification and host prediction, demonstrating promising potential for genomic understanding. However, they exhibit significant limitations in complex reasoning tasks involving long-range dependencies and fine-grained functional localization. These findings highlight the necessity of developing next-generation models with enhanced reasoning capabilities for biological sequences.

基因组分析大模型噬菌体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。