综述基于Transformer的模型在核酸序列分析中的应用进展。
A Review on the Applications of Transformer-based language models for Nucleotide Sequence Analysis
- 梳理Transformer模型在生物序列分析中的关键应用与适配方法。
- 总结其在基因组注释、突变预测等任务中的有效表现。
- 适合初学者快速掌握模型原理及生物信息学应用方向。
近年来,基于Transformer的语言模型在自然语言处理领域影响深远。由于生物序列与自然语言存在相似性,NLP中的模型可被轻松拓展并应用于生物信息学。本文回顾了过去几年中基于Transformer模型在核酸序列分析方面的重大进展,系统分析了大量相关应用论文,揭示了其核心特征与多种定制化方法。同时,提供了对Transformer工作机制的结构化说明,帮助初次使用者理解复杂架构的本质。本文旨在促进科学界对基于Transformer模型在核酸序列分析中应用的理解,并激励读者在此基础上解决更多生物信息学问题。
原文摘要 · Abstract (English)
In recent times, Transformer-based language models are making quite an impact in the field of natural language processing. As relevant parallels can be drawn between biological sequences and natural languages, the models used in NLP can be easily extended and adapted for various applications in bioinformatics. In this regard, this paper introduces the major developments of Transformer-based models in the recent past in the context of nucleotide sequences. We have reviewed and analysed a large number of application-based papers on this subject, giving evidence of the main characterizing features and to different approaches that may be adopted to customize such powerful computational machines. We have also provided a structured description of the functioning of Transformers, that may enable even first time users to grab the essence of such complex architectures. We believe this review will help the scientific community in understanding the various applications of Transformer-based language models to nucleotide sequences. This work will motivate the readers to build on these methodologies to tackle also various other problems in the field of bioinformatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。