arXiv:2412.13716q-bio.GNcs.LG2024-12NeurIPS被引 21

模型自己学着分词,让DNA序列理解更准更快。

Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA

  • 模型用自适应分词,自动学出适合基因序列的切分方式。
  • 在少量预训练下表现超越现有方法,性能更优。
  • 适合对基因语言建模感兴趣的科研人员和开发者。

基础模型在理解DNA基因组语言方面取得了显著进展。然而,以往模型通常采用自然语言的分词方法,这不适用于具有独特特性的DNA序列。此外,最佳的DNA分词方式尚未被充分探索,即使被发现,也可能难以被人类直观理解。为解决这些问题,我们提出MxDNA,一种新框架,让模型通过梯度下降自主学习有效的DNA分词策略。MxDNA采用稀疏卷积专家混合与可变形卷积来建模分词过程,明确考虑了有意义基因片段的不连续性、重叠性和模糊性。在核苷酸变压器基准和基因组基准上,MxDNA以更少的预训练数据和时间,展现出优于现有方法的性能,证明其有效性。最后,我们发现MxDNA学到的分词策略与以往方法显著不同,并在自监督预训练中捕捉到基因功能的分词级特征。MxDNA为DNA分词提供了新视角,可能在多个领域带来广泛应用并产生深刻见解。

原文摘要 · Abstract (English)

Foundation models have made significant strides in understanding the genomic language of DNA sequences. However, previous models typically adopt the tokenization methods designed for natural language, which are unsuitable for DNA sequences due to their unique characteristics. In addition, the optimal approach to tokenize DNA remains largely under-explored, and may not be intuitively understood by humans even if discovered. To address these challenges, we introduce MxDNA, a novel framework where the model autonomously learns an effective DNA tokenization strategy through gradient decent. MxDNA employs a sparse Mixture of Convolution Experts coupled with a deformable convolution to model the tokenization process, with the discontinuous, overlapping, and ambiguous nature of meaningful genomic segments explicitly considered. On Nucleotide Transformer Benchmarks and Genomic Benchmarks, MxDNA demonstrates superior performance to existing methods with less pretraining data and time, highlighting its effectiveness. Finally, we show that MxDNA learns unique tokenization strategy distinct to those of previous methods and captures genomic functionalities at a token level during self-supervised pretraining. Our MxDNA aims to provide a new perspective on DNA tokenization, potentially offering broad applications in various domains and yielding profound insights.

基因建模自适应分词深度学习DNA序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。