arXiv:2509.25274q-bio.GNcs.AI2025-09被引 2

用DNA序列直接识别结直肠癌增强子,性能优于传统方法。

DNABERT-2: Fine-Tuning a Genomic Language Model for Colorectal Gene Enhancer Classification

  • 基于DNABERT-2的Transformer模型,通过字节对编码处理基因序列。
  • 在234万条序列上训练,达到F1值0.704,召回率达0.835。
  • 首次将二代基因组语言模型用于结直肠癌增强子分类,适合癌症基因组研究者。

基因增强子调控基因开启的时间与位置,但其序列多样性和组织特异性使其在结直肠癌中难以定位。我们采用仅依赖序列的方法,微调了基于字节对编码(BPE)的DNABERT-2模型,该模型可学习可变长度的DNA token。利用贝尔法斯特女王大学约翰斯顿癌症研究中心整理的实验数据,我们构建了一个包含234万条1 kb增强子序列的平衡语料库,采用峰中心提取并进行严格去重(包括反向互补合并),按类别分层划分数据集。使用4096词表大小和232个标记的上下文,经Optuna优化超参数后训练了117M参数的DNABERT-2分类器,并在350,742条保留序列上评估。模型取得PR-AUC 0.759、ROC-AUC 0.743,F1最高为0.704(阈值0.359),召回率0.835,精确率0.609。相比在同一数据集上训练的基于CNN的EnhancerNet,DNABERT-2展现出更强的阈值无关排序能力与更高召回率,尽管点精度略低。据我们所知,这是首个将第二代基因组语言模型与BPE分词应用于结直肠癌增强子分类的研究,证明了仅从DNA序列中捕捉肿瘤相关调控信号的可行性。结果表明,基于Transformer的基因组模型可超越基序级编码,实现调控元件的整体分类,为癌症基因组学提供新路径。下一步将提升精确率,探索混合CNN-Transformer结构,并在独立数据集上验证以增强实际应用价值。

原文摘要 · Abstract (English)

Gene enhancers control when and where genes switch on, yet their sequence diversity and tissue specificity make them hard to pinpoint in colorectal cancer. We take a sequence-only route and fine-tune DNABERT-2, a transformer genomic language model that uses byte-pair encoding to learn variable-length tokens from DNA. Using assays curated via the Johnston Cancer Research Centre at Queen's University Belfast, we assembled a balanced corpus of 2.34 million 1 kb enhancer sequences, applied summit-centered extraction and rigorous de-duplication including reverse-complement collapse, and split the data stratified by class. With a 4096-term vocabulary and a 232-token context chosen empirically, the DNABERT-2-117M classifier was trained with Optuna-tuned hyperparameters and evaluated on 350742 held-out sequences. The model reached PR-AUC 0.759, ROC-AUC 0.743, and best F1 0.704 at an optimized threshold (0.359), with recall 0.835 and precision 0.609. Against a CNN-based EnhancerNet trained on the same data, DNABERT-2 delivered stronger threshold-independent ranking and higher recall, although point accuracy was lower. To our knowledge, this is the first study to apply a second-generation genomic language model with BPE tokenization to enhancer classification in colorectal cancer, demonstrating the feasibility of capturing tumor-associated regulatory signals directly from DNA sequence alone. Overall, our results show that transformer-based genomic models can move beyond motif-level encodings toward holistic classification of regulatory elements, offering a novel path for cancer genomics. Next steps will focus on improving precision, exploring hybrid CNN-transformer designs, and validating across independent datasets to strengthen real-world utility.

基因组增强子Transformer癌症研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。