arXiv:2505.08918q-bio.GNcs.AI2025-05被引 1

BPE分词发现灵长类基因组重复序列主导词汇表,影响比较基因组分析

When repeats drive the vocabulary: a Byte-Pair Encoding analysis of T2T primate genomes

  • 用自研dnaBPE对9个端到端基因组做独立分词,固定512000词表
  • 仅11569个词共享所有基因组,近99万词仅属单一基因组
  • 重复元件过量干扰分词结果,提示需结合重复序列屏蔽优化

端到端(T2T)基因组组装的出现为比较基因组学开辟了新途径,但针对基因组序列的有效分词策略仍研究不足。本初步研究使用字节对编码(BPE)对九个T2T灵长类基因组(包括三个人类组装)进行分析,采用自研工具dnaBPE,训练独立的BPE分词器,固定词表大小为512,000。结果显示,仅有11,569个词在所有基因组中共享,而近991,854个词仅属于单一基因组,表明随着基因组比较数量增加,共享词汇迅速减少。此外,基于词重叠构建的系统发育树未能复现已知的灵长类亲缘关系,这归因于物种特异性高拷贝重复元件的显著影响。这些发现揭示了BPE分词的双重性:虽能有效压缩重复序列,但对高拷贝元件高度敏感,限制其作为通用比较基因组学工具的应用。我们讨论了混合策略与重复序列屏蔽方法以改进基因组分词,强调开发大规模基因组语言模型需进行领域特定适配。本研究使用的dnaBPE工具开源,可于https://github.com/aglabx/dnaBPE 获取。

原文摘要 · Abstract (English)

The emergence of telomere-to-telomere (T2T) genome assemblies has opened new avenues for comparative genomics, yet effective tokenization strategies for genomic sequences remain underexplored. In this pilot study, we apply Byte Pair Encoding (BPE) to nine T2T primate genomes including three human assemblies by training independent BPE tokenizers with a fixed vocabulary of 512,000 tokens using our custom tool, dnaBPE. Our analysis reveals that only 11,569 tokens are shared across all assemblies, while nearly 991,854 tokens are unique to a single genome, indicating a rapid decline in shared vocabulary with increasing assembly comparisons. Moreover, phylogenetic trees derived from token overlap failed to recapitulate established primate relationships, a discrepancy attributed to the disproportionate influence of species-specific high-copy repetitive elements. These findings underscore the dual nature of BPE tokenization: while it effectively compresses repetitive sequences, its sensitivity to high-copy elements limits its utility as a universal tool for comparative genomics. We discuss potential hybrid strategies and repeat-masking approaches to refine genomic tokenization, emphasizing the need for domain-specific adaptations in the development of large-scale genomic language models. The dnaBPE tool used in this study is open-source and available at https://github.com/aglabx/dnaBPE.

基因组分析BPE分词重复序列语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。