arXiv:2410.06211q-bio.GNcs.LG2024-10被引 5

提出可直接读取基因调控机制的神经网络,让模型发现的规律真正看得懂。

A mechanistically interpretable neural network for regulatory genomics

  • 设计新架构,让调控基序和语法直接体现在权重与激活中
  • 在新序列生成、基序发现上表现优异,且对上下文变化不敏感
  • 适合需要理解基因调控原理的研究者,如功能基因组学专家

深度神经网络在将基因组DNA序列映射到相关读数(如蛋白质-DNA结合)方面表现出色。除了预测能力,这些模型还应揭示驱动基因调控的底层基序及其语法结构。传统方法从卷积滤波器中提取基序时,信息分散于多层滤波器中,难以解释;依赖重要性得分的方法又常不稳定不可靠。为此,我们设计了一种新型可机制解释的基因组学神经网络架构,使基序及其语法能直接从学习到的权重和激活中读取。我们提供了该架构全表达性的理论与实证证据,同时保持高度可解释性。通过多项实验,证明该架构在从头基序发现、基序实例识别上表现卓越,对序列上下文变化具有鲁棒性,并能实现完全可解释的新功能序列生成。

原文摘要 · Abstract (English)

Deep neural networks excel in mapping genomic DNA sequences to associated readouts (e.g., protein-DNA binding). Beyond prediction, the goal of these networks is to reveal to scientists the underlying motifs (and their syntax) which drive genome regulation. Traditional methods that extract motifs from convolutional filters suffer from the uninterpretable dispersion of information across filters and layers. Other methods which rely on importance scores can be unstable and unreliable. Instead, we designed a novel mechanistically interpretable architecture for regulatory genomics, where motifs and their syntax are directly encoded and readable from the learned weights and activations. We provide theoretical and empirical evidence of our architecture's full expressivity, while still being highly interpretable. Through several experiments, we show that our architecture excels in de novo motif discovery and motif instance calling, is robust to variable sequence contexts, and enables fully interpretable generation of novel functional sequences.

可解释性基因调控神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。