Wisteria统一建模DNA序列的局部与全局特征,提升基因组语言理解能力。
Wisteria: A Unified Multi-Scale Feature Learning Framework for DNA Language Model

- 融合门控膨胀卷积与多层感知机,同时捕捉局部基序与长程依赖
- 在四种设置下优于现有模型,长序列任务表现突出
- 适合基因组功能预测、序列分析等生物信息学研究者使用
DNA语言模型旨在通过捕捉DNA序列中的长距离依赖关系,解析基因组的调控语法与语义。现有方法虽重视长程标记交互,却常忽略局部基序与全局依赖之间的相互作用。本文提出Wisteria,一种在统一框架内实现多尺度特征学习的基因组语言模型。具体而言,Wisteria在基于Mamba的架构上引入门控膨胀卷积以捕获局部基序和调控模式,同时使用门控多层感知机优化全局依赖。此外,我们设计了一种基于傅里叶的注意力机制,支持频域建模、周期扩展与长度泛化。在包含短距与长距依赖的四个实验设置中,Wisteria在下游基准测试中均表现出色,优于多种竞争性DNA语言模型。结果表明,Wisteria能有效统一局部与全局依赖建模,适用于多尺度基因组序列分析。
原文摘要 · Abstract (English)
DNA language model aims to decipher the regulatory grammar and semantic of genomes by capturing long range dependencies in DNA sequences. Existing methods emphasize long range token interactions but often ignore the interplay between local motifs and global dependencies. In this paper, we propose Wisteria, a genomic language model that integrates multi scale feature learning within a unified framework for DNA sequence. Specifically, Wisteria augments the Mamba based architecture with gated dilated convolutions to capture local motifs and regulatory patterns, while gated multilayer perceptrons refine global dependencies. We further introduce a Fourier based attention mechanism to support frequency domain modeling, periodic extension and length generalization. Across four experimental settings with both short and long range dependencies, Wisteria demonstrates strong performance on downstream benchmarks against competitive DNA language model baselines. These results indicate that Wisteria effectively unifies local and global dependency modeling for multi scale genomic sequence analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。