用深度学习提升重复序列分类准确率,覆盖更广物种。
Terrier: A Deep Learning Repeat Classifier
- 基于深度学习模型Terrier,利用Repbase数据库训练分类重复序列。
- 在人类等模式生物中分类准确率超越现有方法,覆盖97.1%的重复序列。
- 适用于非模式生物,助力研究基因组演化与表型变异。
重复DNA序列是基因组结构与进化过程的基础,但其准确分类仍具挑战。Terrier是一种深度学习模型,利用公开且经整理的重复序列库(遵循RepeatMasker标准)进行训练,以克服现有方法因物种代表性不足导致的分类误差与可复现性差的问题。模型在包含超过10万种重复家族的Repbase数据集上训练,远超Dfam的数据量,实现了对97.1% Repbase序列的准确映射至RepeatMasker类别,构建了当前最全面的分类体系。在水稻、果蝇、人类和小鼠等模式生物中,相比DeepTE、TERL和TEclass2,Terrier展现出更高精度并能处理更广泛的序列。在非模式生物(两栖类、扁虫、北方磷虾)中的验证进一步证明其有效性,有助于推动重复序列驱动的基因组演化、基因组不稳定性及表型变异研究。
原文摘要 · Abstract (English)
Repetitive DNA sequences underpin genome architecture and evolutionary processes, yet they remain challenging to classify accurately. Terrier is a deep learning model designed to overcome these challenges by classifying repetitive DNA sequences using a publicly available, curated repeat sequence library trained under the RepeatMasker schema. Poor representation of taxa within repeat databases often limits the classification accuracy and reproducibility of current repeat annotation methods, limiting our understanding of repeat evolution and function. Terrier overcomes these challenges by leveraging deep learning for improved accuracy. Trained on Repbase, which includes over 100,000 repeat families -- four times more than Dfam -- Terrier maps 97.1% of Repbase sequences to RepeatMasker categories, offering the most comprehensive classification system available. When benchmarked against DeepTE, TERL, and TEclass2 in model organisms (rice, fruit flies, humans, and mice), Terrier achieved superior accuracy while classifying a broader range of sequences. Further validation in non-model amphibian, flatworm and Northern krill genomes highlights its effectiveness in improving classification in non-model species, facilitating research on repeat-driven evolution, genomic instability, and phenotypic variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。