用深度学习识别基因敲除前后染色质开放区的差异序列特征
WTKO-CNN: Deep Learning Reveals Sequence Motifs Distinguishing Wild-Type and Knockout ATAC-seq Peaks

- 基于注意力机制的CNN模型区分野生型与敲除型染色质峰
- 通过显著性图发现关键核苷酸位置并提取新发基序
- 结果与已知转录因子结合位点高度一致,适合基因调控研究者
染色质调节因子可通过改变调控DNA元件的可及性来影响转录程序。理解野生型(WT)与敲除型(KO)条件下调控序列的差异,对解析转录控制至关重要。本文提出一种带注意力机制的卷积神经网络(WTKO-CNN),用于分类DNA序列是否为WT或KO,表现优异。为解释模型决策,生成显著性图以定位对分类最具影响力的核苷酸位置。从高显著性区域提取并聚类k-mer,实现无监督基序发现。由CNN滤波器生成的序列图谱和共识基序揭示了具有生物学意义的模式,并通过MEME、TOMTOM和HOMER与已知转录因子结合位点进行验证。分析识别出能区分WT与KO序列的转录因子家族相关基序,证明基于CNN显著性映射是发现功能序列特征的有效方法。
原文摘要 · Abstract (English)
Chromatin regulators can alter transcriptional programs by modifying the accessibility of regulatory DNA elements. Understanding how regulatory sequences differ between wild-type (WT) and knockout (KO) conditions is crucial for deciphering transcriptional control. Here, we applied a convolutional neural network, \textbf{WTKO-CNN} with an attention mechanism to classify DNA sequences as WT or KO, achieving high predictive performance. To interpret the model, we generated saliency maps to identify nucleotide positions most influential for the classification decision. From these high-saliency regions, we extracted and clustered k-mers, enabling de novo motif discovery. Sequence logos and consensus motifs derived from the CNN filters revealed biologically meaningful patterns, which are further validated using MEME, TOMTOM, and HOMER against known transcription factor binding sites. Our analysis identified motifs associated with transcription factor families that discriminate WT from KO sequences, demonstrating that CNN-guided saliency mapping is a powerful approach for uncovering functional sequence features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。