arXiv:2602.21550cs.LGq-bio.GN2026-02中稿 · ICLR被引 1

用多模态表观基因组信号提升基因表达预测,短序列也能超越长序列模型。

Extending Sequence Length is Not All You Need: Effective Integration of Multimodal Signals for Gene Expression Prediction

  • 设计Prism框架,学习多种表观遗传特征组合以表征染色质背景状态。
  • 在短序列下实现最优性能,证明近端信号比远端序列更重要。
  • 通过反向门调整缓解背景模式带来的混淆,适合生物信息学研究者使用。

基因表达预测旨在从DNA序列预测mRNA表达水平,面临巨大挑战。以往工作聚焦于扩展输入序列长度以定位远端增强子(可影响目标基因数百千碱基外)。我们首次发现,当前模型中长序列建模反而会降低性能,即便精心设计算法也只能缓解此问题。相反,我们发现目标基因附近的近端多模态表观基因组信号更为关键。不同信号类型具有不同生物学作用:部分直接标记活跃调控元件,另一些反映背景染色质状态,可能引入混淆效应。简单拼接会导致模型错误关联这些背景模式。为此,我们提出Prism框架,学习高维表观基因组特征的多种组合以表示不同染色质背景状态,并采用反向门调整方法缓解混淆。实验表明,合理建模多模态信号可在仅使用短序列时达到最先进性能。

原文摘要 · Abstract (English)

Gene expression prediction, which predicts mRNA expression levels from DNA sequences, presents significant challenges. Previous works often focus on extending input sequence length to locate distal enhancers, which may influence target genes from hundreds of kilobases away. Our work first reveals that for current models, long sequence modeling can decrease performance. Even carefully designed algorithms only mitigate the performance degradation caused by long sequences. Instead, we find that proximal multimodal epigenomic signals near target genes prove more essential. Hence we focus on how to better integrate these signals, which has been overlooked. We find that different signal types serve distinct biological roles, with some directly marking active regulatory elements while others reflect background chromatin patterns that may introduce confounding effects. Simple concatenation may lead models to develop spurious associations with these background patterns. To address this challenge, we propose Prism, a framework that learns multiple combinations of high-dimensional epigenomic features to represent distinct background chromatin states and uses backdoor adjustment to mitigate confounding effects. Our experimental results demonstrate that proper modeling of multimodal epigenomic signals achieves state-of-the-art performance using only short sequences for gene expression prediction.

基因表达多模态融合表观基因组深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。