arXiv:2507.08542cs.LG2025-07

用深度学习直接从植物基因组预测环状RNA剪接位点,速度快且能发现新环状RNA。

CircFormerMoE: An End-to-End Deep Learning Framework for Circular RNA Splice Site Detection and Pairing in Plant Genomes

  • 基于Transformer和专家混合模型,直接从基因组序列预测剪接位点
  • 在10种植物中验证,可发现未标注的环状RNA,准确率高
  • 适用于大规模植物环状RNA发现,对功能基因组研究有重要意义

环状RNA(circRNA)是非编码RNA调控网络的重要组成部分。以往的识别方法主要依赖高通量RNA测序数据结合比对算法检测反向剪接信号,但存在诸多局限:无法直接从基因组DNA序列预测circRNA,严重依赖实验数据;计算成本高,因复杂比对与过滤步骤;在大规模或全基因组预测中效率低下。植物中问题更严峻,因植物circRNA剪接位点常缺乏人类mRNA中常见的GT-AG保守序列,且尚无具备强泛化能力的高效深度学习模型。此外,目前鉴定的植物circRNA数量远低于真实丰度。本文提出一种名为CircFormerMoE的端到端深度学习框架,基于Transformer与专家混合模型,直接从植物基因组DNA预测circRNA。框架包含剪接位点检测(SSD)和剪接位点配对(SSP)两个子任务。在10种植物基因数据上验证了其有效性,训练于已知circRNA实例后,仍可发现未标注的新circRNA。此外,通过可解释性分析揭示了影响预测的关键序列模式。该框架为植物中大规模circRNA发现提供了一种快速、准确的计算方法与工具,为未来植物功能基因组学与非编码RNA注释研究奠定基础。

原文摘要 · Abstract (English)

Circular RNAs (circRNAs) are important components of the non-coding RNA regulatory network. Previous circRNA identification primarily relies on high-throughput RNA sequencing (RNA-seq) data combined with alignment-based algorithms that detect back-splicing signals. However, these methods face several limitations: they can't predict circRNAs directly from genomic DNA sequences and relies heavily on RNA experimental data; they involve high computational costs due to complex alignment and filtering steps; and they are inefficient for large-scale or genome-wide circRNA prediction. The challenge is even greater in plants, where plant circRNA splice sites often lack the canonical GT-AG motif seen in human mRNA splicing, and no efficient deep learning model with strong generalization capability currently exists. Furthermore, the number of currently identified plant circRNAs is likely far lower than their true abundance. In this paper, we propose a deep learning framework named CircFormerMoE based on transformers and mixture-of experts for predicting circRNAs directly from plant genomic DNA. Our framework consists of two subtasks known as splicing site detection (SSD) and splicing site pairing (SSP). The model's effectiveness has been validated on gene data of 10 plant species. Trained on known circRNA instances, it is also capable of discovering previously unannotated circRNAs. In addition, we performed interpretability analyses on the trained model to investigate the sequence patterns contributing to its predictions. Our framework provides a fast and accurate computational method and tool for large-scale circRNA discovery in plants, laying a foundation for future research in plant functional genomics and non-coding RNA annotation.

环状RNA深度学习植物基因组剪接预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。