CNN在DNA基础模型中重获优势,性能超越多数Transformer与SSM方法。
Revisiting Convolution Architecture in the Realm of DNA Foundation Models
- 提出ConvNova:融合空洞卷积、门控机制与双分支结构的CNN设计。
- 在超过一半任务上超越现有方法,组蛋白相关任务平均领先5.8%。
- 参数更少、计算更快,适合生物序列建模场景,激发对CNN新关注。
近年来,基于Transformer和状态空间模型(SSM)的方法推动了DNA基础语言模型的发展。然而,这些新方法与经典卷积网络(CNN)在基础模型基准上的对比仍显不足。本文提出一种简洁而精心设计的CNN方法ConvNova,包含三个有效设计:空洞卷积、门控卷积及双分支门控框架。大量实验证明,ConvNova在多个基础模型基准的超半数任务上显著优于现有方法。例如,在组蛋白相关任务中,其平均性能超过次优方法5.8%,且通常使用更少参数并实现更快计算。实验还发现与生物特性相关的现象,表明CNN在当前仍具备强大竞争力。本工作有望重新引发对基于CNN的DNA基础模型研究的兴趣。
原文摘要 · Abstract (English)
In recent years, a variety of methods based on Transformer and state space model (SSM) architectures have been proposed, advancing foundational DNA language models. However, there is a lack of comparison between these recent approaches and the classical architecture convolutional networks (CNNs) on foundation model benchmarks. This raises the question: are CNNs truly being surpassed by these recent approaches based on transformer and SSM architectures? In this paper, we develop a simple but well-designed CNN-based method termed ConvNova. ConvNova identifies and proposes three effective designs: 1) dilated convolutions, 2) gated convolutions, and 3) a dual-branch framework for gating mechanisms. Through extensive empirical experiments, we demonstrate that ConvNova significantly outperforms recent methods on more than half of the tasks across several foundation model benchmarks. For example, in histone-related tasks, ConvNova exceeds the second-best method by an average of 5.8%, while generally utilizing fewer parameters and enabling faster computation. In addition, the experiments observed findings that may be related to biological characteristics. This indicates that CNNs are still a strong competitor compared to Transformers and SSMs. We anticipate that this work will spark renewed interest in CNN-based methods for DNA foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。