测试大模型直接理解基因组序列的能力,发现其能识别局部信号但难处理复杂推理。
GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
- 设计新基准GenomeQA,用5200个真实基因序列评估通用大模型
- 模型在短序列特征识别上表现良好,但长程模式推理能力不足
- 适合研究基因组大模型的开发者和生物信息学交叉学者
大语言模型(LLMs)在基因组学中被广泛用作自然语言接口的智能助手,用于推理生物学知识、注释和分析结果。然而,现有基准要么聚焦于专门训练的DNA预测模型,要么仅通过文本问题评估生物知识,未充分探索通用大模型直接处理原始基因组序列的行为。我们提出GenomeQA,一个针对通用大模型在序列基因组推理任务上的受控评估基准。该基准包含从多个生物数据库抽取的5200个样本,序列长度为6至1000碱基对,涵盖六类任务:增强子与启动子识别、剪接位点识别、分类学分类、组蛋白修饰预测、转录因子结合位点预测及TF基序预测。在六种前沿大模型上测试发现,模型普遍优于随机基线,可利用局部序列信号如GC含量和短基序,但在需要间接或多步推理的任务上性能下降。GenomeQA为研究和改进通用大模型在原始基因组序列上的应用提供了诊断性基准。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly adopted as conversational assistants in genomics, where they are mainly used to reason over biological knowledge, annotations, and analysis outputs through natural language interfaces. However, existing benchmarks either focus on specialized DNA models trained for sequence prediction or evaluate biological knowledge using text-only questions, leaving the behavior of general-purpose LLMs when directly exposed to raw genome sequences underexplored. We introduce GenomeQA, a benchmark designed to provide a controlled evaluation setting for general-purpose LLMs on sequence-based genome inference tasks. GenomeQA comprises 5,200 samples drawn from multiple biological databases, with sequence lengths ranging from 6 to 1,000 base pairs (bp), spanning six task families: Enhancer and Promoter Identification, Splice Site Identification, Taxonomic Classification, Histone Mark Prediction, Transcription Factor Binding Site Prediction, and TF Motif Prediction. Across six frontier LLMs, we find that models consistently outperform random baselines and can exploit local sequence signals such as GC content and short motifs, while performance degrades on tasks that require more indirect or multi-step inference over sequence patterns. GenomeQA establishes a diagnostic benchmark for studying and improving the use of general-purpose LLMs on raw genomic sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。