arXiv:2604.16570cs.LGcs.AI2026-04

揭示DNA预训练三大被忽视问题,提出可复现的评估标准

In Search of Lost DNA Sequence Pretraining

论文配图:In Search of Lost DNA Sequence Pretraining
图 1 · 摘自论文原文
  • 指出当前预训练忽略评估数据集选择、掩码策略缺陷和词表设计
  • 实验验证问题严重性,发现错误设置导致性能偏差超30%
  • 提供标准化测试平台,助力基因组基础模型严谨对比

DNA序列编码是基因功能预测、蛋白质合成及多种下游生物任务的基础。尽管大规模DNA序列预训练取得显著进展,现有研究过度关注预训练规模和定制化下游数据集,而忽视了预训练范式中一些关键环节。本文揭示了三个长期被忽视的核心问题:下游数据集选择不当、邻域掩码策略存在固有缺陷,以及词汇表讨论不足。为此,我们开展了全面调查,提出了系统性指导原则,包括评估数据集的选择标准、任务设计指南和深度词表分析方法。大量实验证明了这些问题的重要性,并验证了建议的合理性。最后,我们构建了一个标准化测试平台,支持可复现、严谨的DNA预训练方法基准测试,推动基因组基础模型的发展。

原文摘要 · Abstract (English)

DNA sequence encoding is fundamental to gene function prediction, protein synthesis, and diverse downstream biological tasks. Despite the substantial progress achieved by large-scale DNA sequence pretraining, existing studies have overwhelmingly emphasized pretraining scale and custom downstream evaluation datasets, while neglecting some essential components of the pretraining paradigm. In this paper, we reveal three critical yet heretofore overlooked problems in DNA pretraining: inappropriate downstream datasets, inherent flaws in the neighbor-masking strategy, and the lack of detailed discussion on vocabulary. Therefore, we undertake comprehensive investigations and propose principled guidelines, including selection criteria for evaluation datasets, guiding task design, and in-depth vocabulary analysis. Extensive experiments validate the significance of our identified problems and support the rationale behind our recommendations. Finally, we introduce a standardized testbed that enables reproducible and rigorous benchmarking of DNA pretraining methods to advance the development of genomic foundation models.

基因组模型预训练评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。