arXiv:2606.30140q-bio.GNcs.CL2026-06

对比Transformer与传统模型在基因序列任务中的预训练效果

DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks

论文配图:DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks
图 1 · 摘自论文原文
  • 采用Transformer与卷积模型对比预训练性能
  • 发现预训练对下游任务提升有限,非必需
  • 质疑BPE分词在基因序列中的适用性

近期大模型突破为基因组序列研究带来新机遇。以DNABERT2为代表的Transformer模型与以ConvNova为代表的卷积模型并存,但系统性对比仍匮乏。鉴于Transformer需大量昂贵预训练,其性能增益是否值得成为关键问题。此外,如DNABERT2等模型依赖字节对编码(BPE)分词,其在基因序列表征中的有效性尚存争议。本文围绕三个核心问题展开:(i) Transformer模型在重预训练后能否显著提升微调任务表现?(ii) 预训练在该场景中实际贡献几何?(iii) BPE分词对基因组相关任务性能有何影响?

原文摘要 · Abstract (English)

Recent breakthroughs in foundation models and Large Language Models (LLMs) have introduced new opportunities for studying and decoding genomic sequences. Several state-of-the-art approaches, such as DNABERT2, rely on transformer-based architectures, while others, such as ConvNova, still build upon more conventional convolutional models. However, systematic benchmark comparisons across these methods remain scarce. Given that transformer-based models require extensive and costly pretraining, it is crucial to evaluate whether their performance gains justify this overhead. Moreover, LLMs such as DNABERT2 typically rely on Byte Pair Encoding (BPE) tokenization, whose relevance for DNA sequence representation is still debated within the genomics community. In this work, we investigate three key questions: (i) do transformer-based models provide sufficient improvements on fine-tuning tasks upon heavy pretraining, (ii) what is the actual contribution of pretraining in this setting, and (iii) how does BPE tokenization impact performance on genomics-related tasks?

基因组学预训练Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。