arXiv:2502.13991q-bio.GNcs.AI2025-02ICLR被引 5

通过因果建模发现调控基因表达的关键序列元件,提升预测精度。

Learning to Discover Regulatory Elements for Gene Expression Prediction

论文配图:Learning to Discover Regulatory Elements for Gene Expression Prediction
图 1 · 摘自论文原文
  • 基于因果调控元件分解表观基因组信号与DNA序列
  • 在多个数据集上优于现有基线模型,准确率显著提升
  • 适合基因调控研究者和生物信息学方向的开发者

我们研究从DNA序列预测基因表达的问题。该任务的核心挑战在于识别控制基因表达的调控元件。为此,我们提出Seq2Exp,一种专门设计用于发现并提取驱动目标基因表达的调控元件的序列到表达网络,从而提升基因表达预测的准确性。该方法捕捉了表观基因组信号、DNA序列与其相关调控元件之间的因果关系。具体而言,我们提出在因果活跃调控元件条件下分解表观基因组信号和DNA序列,并利用基于β分布的信息瓶颈机制整合其影响,同时过滤非因果成分。实验表明,Seq2Exp在基因表达预测任务中优于现有基线模型,并能发现比MACS3等常用峰值检测方法更具影响力的区域。源代码已作为AIRS库的一部分发布(https://github.com/divelab/AIRS/)。

原文摘要 · Abstract (English)

We consider the problem of predicting gene expressions from DNA sequences. A key challenge of this task is to find the regulatory elements that control gene expressions. Here, we introduce Seq2Exp, a Sequence to Expression network explicitly designed to discover and extract regulatory elements that drive target gene expression, enhancing the accuracy of the gene expression prediction. Our approach captures the causal relationship between epigenomic signals, DNA sequences and their associated regulatory elements. Specifically, we propose to decompose the epigenomic signals and the DNA sequence conditioned on the causal active regulatory elements, and apply an information bottleneck with the Beta distribution to combine their effects while filtering out non-causal components. Our experiments demonstrate that Seq2Exp outperforms existing baselines in gene expression prediction tasks and discovers influential regions compared to commonly used statistical methods for peak detection such as MACS3. The source code is released as part of the AIRS library (https://github.com/divelab/AIRS/).

基因表达调控元件因果建模深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。