arXiv:2506.02212cs.CLcs.AI2025-06综述被引 3

用语言模型分析基因序列,揭示生命奥秘

Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics

  • 将NLP中的词向量与变换器技术用于解析DNA、RNA和蛋白质序列
  • 通过分词策略与模型架构优化,提升对基因组数据的解读能力
  • 适合生物信息学研究者探索基因功能与进化关系

自然语言处理(NLP)已突破语言领域,被应用于生物序列分析。本文综述NLP在基因组学、转录组学和蛋白质组学中的应用,涵盖从word2vec到基于Transformer与海豹算子的先进模型。这些方法被用于解析DNA、RNA、蛋白质序列及全基因组数据,重点探讨分词策略与模型结构对不同生物任务的适配性。同时介绍近期进展,包括蛋白质结构预测、基因表达分析与进化研究,展示其从大规模基因组数据中提取关键信息的潜力。随着语言模型持续演进,其与生物信息学的融合将极大推动对生命过程的理解。

原文摘要 · Abstract (English)

Natural Language Processing (NLP) has transformed various fields beyond linguistics by applying techniques originally developed for human language to the analysis of biological sequences. This review explores the application of NLP methods to biological sequence data, focusing on genomics, transcriptomics, and proteomics. We examine how various NLP methods, from classic approaches like word2vec to advanced models employing transformers and hyena operators, are being adapted to analyze DNA, RNA, protein sequences, and entire genomes. The review also examines tokenization strategies and model architectures, evaluating their strengths, limitations, and suitability for different biological tasks. We further cover recent advances in NLP applications for biological data, such as structure prediction, gene expression, and evolutionary analysis, highlighting the potential of these methods for extracting meaningful insights from large-scale genomic data. As language models continue to advance, their integration into bioinformatics holds immense promise for advancing our understanding of biological processes in all domains of life.

NLP基因组学生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。