RIBOSPAN可精准建模长达1万碱基的完整RNA,突破传统模型长度限制。
RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling

- 采用单碱基分词与注意力隔离序列打包,支持超长上下文建模。
- 在10,240碱基下仍保持强重建能力与局部化扰动响应。
- 适用于长链RNA分析、突变预测及全转录本mRNA设计,适合生物研究者。
全长RNA(尤其是信使RNA)常超过现有RNA基础模型的上下文长度,限制了单碱基分辨率的完整转录本建模。我们提出RIBOSPAN,一个16.1亿参数的双向RNA基础模型,原生预训练上下文长度可达10,240碱基。RIBOSPAN结合密集双向自注意力、单碱基分词和注意力隔离序列打包,实现对完整长RNA的高分辨率建模。原生10K预训练在10,240令牌下仍保持强重构能力;在受控长上下文基准测试中,维持强上下文响应性与上下文特异性表征分离,同时扰动引起的改变高度局部化。推理时采用YaRN缩放恢复了直接短上下文外推丢失的上下文组织,但导致更显著的远距离表征扩散。冻结类型评估显示,RIBOSPAN在长RNA上学习到最先进表示,尤其在长序列上优势明显。在下游生物基准测试中,RIBOSPAN成为最强的编码器类RNA基础模型,在全转录本生物属性预测与零样本突变适应度建模中均达到最先进水平。基于相同骨干,我们构建了多维条件离散扩散框架,用于全长mRNA生成与重设计,包括保持蛋白质不变的同义密码子扩散优化。RIBOSPAN为可迁移的RNA表征学习、生物预测与全转录本mRNA设计提供了强大基础。
原文摘要 · Abstract (English)
Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution. We present RIBOSPAN, a 1.61-billion-parameter bidirectional RNA foundation model natively pretrained with context lengths up to 10,240 nt. RIBOSPAN combines dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing to enable high-resolution modeling of complete long RNAs. Native 10K pretraining preserves strong reconstruction at 10,240 tokens and, in a controlled long-context benchmark, maintains strong contextual responsiveness and context-specific representation separation while keeping perturbation-induced changes highly localized. Inference-time YaRN scaling recovers much of the contextual organization lost by direct short-context extrapolation, but induces substantially greater distal representation diffusion. Frozen RNA-type evaluations show that RIBOSPAN learns state-of-the-art RNA representations, with a particularly clear advantage on long RNAs. Across downstream biological benchmarks, RIBOSPAN emerges as the strongest encoder-only RNA foundation model, achieving state-of-the-art performance in both full-transcript biological property prediction and zero-shot mutation-fitness modeling. Building on the same backbone, we develop a multidimensionally conditioned discrete-diffusion framework for full-length mRNA generation and redesign, including synonymous-codon diffusion for protein-preserving CDS optimization. Together, RIBOSPAN establishes a powerful long-context foundation for transferable RNA representation learning, biological prediction, and full-transcript mRNA design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。