dnaGrinder高效处理基因组长程依赖,性能超越主流模型。
dnaGrinder: a lightweight and high-capacity genomic foundation model
- 基于轻量架构,有效建模基因序列长距离依赖
- 在17,000+令牌输入下表现优异,单卡支持超14万令牌序列
- 可在工作站级显卡微调,适合科研与临床应用
理解基因组序列中编码的复杂信息仍是生物研究与临床应用中的重大挑战。近年来,大语言模型的发展催生了专用于解析DNA序列的编码器-仅和解码器-仅基础模型。然而,仍存在诸多问题:基因组序列中长程依赖的高效管理、核苷酸变异的有效表征,以及大模型架构与大规模预训练数据带来的高昂计算成本。现有基因组基础模型常面临性能与规模的权衡——小模型表现一般,大模型性能虽好但资源消耗高。为此,我们提出dnaGrinder,一种独特且高效的基因组基础模型。它在不牺牲性能的前提下,有效处理基因组序列中的长程依赖,并显著降低计算开销。其性能不仅可比,往往优于Nucleotide Transformer和DNABERT-2等领先模型。此外,dnaGrinder设计为可在工作站级GPU上轻松微调,支持超过17,000个标记的输入长度;在单块高性能GPU上,可处理超过140,000个标记的序列,是基础生物学研究与临床应用中高效且易用的工具。
原文摘要 · Abstract (English)
The task of understanding and interpreting the complex information encoded within genomic sequences remains a grand challenge in biological research and clinical applications. In this context, recent advancements in large language model research have led to the development of both encoder-only and decoder-only foundation models designed to decode intricate information in DNA sequences. However, several issues persist, particularly regarding the efficient management of long-range dependencies inherent in genomic sequences, the effective representation of nucleotide variations, and the considerable computational costs associated with large model architectures and extensive pretraining datasets. Current genomic foundation models often face a critical tradeoff: smaller models with mediocre performance versus large models with improved performance. To address these challenges, we introduce dnaGrinder, a unique and efficient genomic foundation model. dnaGrinder excels at managing long-range dependencies within genomic sequences while minimizing computational costs without compromising performance. It achieves results that are not just comparable but often superior to leading DNA models such as Nucleotide Transformer and DNABERT-2. Furthermore, dnaGrinder is designed for easy fine-tuning on workstation-grade GPUs, accommodating input lengths exceeding 17,000 tokens. On a single high-performance GPU, it supports sequences longer than 140,000 tokens, making it a highly efficient and accessible tool for both basic biological research and clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。