arXiv:2503.04413cs.CL2025-03

大语言模型可准确预测耐药基因,且能融合文本信息提升性能。

Can Large Language Models Predict Antimicrobial Resistance Gene?

  • 用生成式大模型处理DNA序列,结合文本信息进行分类。
  • 在耐药基因数据集上表现媲美甚至优于传统模型。
  • 适合对生物序列分析感兴趣的开发者和研究人员。

本研究证明,生成式大语言模型在DNA序列分析与分类任务中可比传统Transformer编码器模型更具灵活性。尽管近期的编码器模型如DNABERT和Nucleotide Transformer在DNA分类任务中表现出色,但基于Transformer解码器的生成式大模型在此领域尚未得到充分探索。本文评估了生成式大语言模型在不同标签的DNA序列上的表现,并分析了加入额外文本信息后的性能变化。实验针对耐药基因展开,结果表明,生成式大语言模型在结合序列与文本信息时,能实现相当或更优的预测效果,展现出良好的灵活性与准确性。相关代码与数据已公开于GitHub:https://github.com/biocomgit/llm4dna。

原文摘要 · Abstract (English)

This study demonstrates that generative large language models can be utilized in a more flexible manner for DNA sequence analysis and classification tasks compared to traditional transformer encoder-based models. While recent encoder-based models such as DNABERT and Nucleotide Transformer have shown significant performance in DNA sequence classification, transformer decoder-based generative models have not yet been extensively explored in this field. This study evaluates how effectively generative Large Language Models handle DNA sequences with various labels and analyzes performance changes when additional textual information is provided. Experiments were conducted on antimicrobial resistance genes, and the results show that generative Large Language Models can offer comparable or potentially better predictions, demonstrating flexibility and accuracy when incorporating both sequence and textual information. The code and data used in this work are available at the following GitHub repository: https://github.com/biocomgit/llm4dna.

大模型耐药基因DNA分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。