用基因组序列直接发现非编码RNA,无需转录组数据。
RNAMunin: A Deep Machine Learning Model for Non-coding RNA Discovery
- 基于基因组序列训练的小型深度学习模型,不依赖转录组数据。
- 在约60Gbp的宏基因组数据上训练,可处理多Gb级组装序列。
- 仅100万参数,速度快,适合大规模微生物基因组分析。
微生物基因组的功能注释常偏向于编码蛋白的基因,大量非编码RNA(ncRNAs)未被探索,而这些ncRNAs对细菌和古菌的生理、应激反应和代谢调节至关重要。从基因组序列中直接识别ncRNAs是生物信息学和生物学的重大挑战,对理解生物体的完整调控潜力至关重要。本文提出RNAMunin,一种仅使用基因组序列即可发现ncRNAs的机器学习模型。该模型在约60 Gbp的长读长宏基因组数据(来自旧金山河口16个样本)中提取的Rfam序列上训练,适用于包含数Gb级连续序列的大规模数据集。与现有大多数模型不同,RNAMunin仅需基因组序列输入,无需转录组数据,因此能发现未被转录的ncRNAs。模型规模小(约100万参数),计算高效,适合大规模应用。本文以叙事风格详细阐述模型开发与工作原理。
原文摘要 · Abstract (English)
Functional annotation of microbial genomes is often biased toward protein-coding genes, leaving a vast, unexplored landscape of non-coding RNAs (ncRNAs) that are critical for regulating bacterial and archaeal physiology, stress response and metabolism. Identifying ncRNAs directly from genomic sequence is a paramount challenge in bioinformatics and biology, essential for understanding the complete regulatory potential of an organism. This paper presents RNAMunin, a machine learning (ML) model that is capable of finding ncRNAs using genomic sequence alone. It is also computationally viable for large sequence datasets such as long read metagenomic assemblies with contigs totaling multiple Gbp. RNAMunin is trained on Rfam sequences extracted from approximately 60 Gbp of long read metagenomes from 16 San Francisco Estuary samples. We know of no other model that can detect ncRNAs based solely on genomic sequence at this scale. Since RNAMunin only requires genomic sequence as input, we do not need for an ncRNA to be transcribed to find it, i.e., we do not need transcriptomics data. We wrote this manuscript in a narrative style in order to best convey how RNAMunin was developed and how it works in detail. Unlike almost all current ML models, at approximately 1M parameters, RNAMunin is very small and very fast.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。