通过挖掘病毒基因组中缺失的序列模式,提升分类准确率。
Mining Negative Sequential Patterns to Improve Viral Genomic Feature Representation and Classification

- 基于基因组中缺失序列模式提取特征,突破传统拼接依赖的局限。
- 在8种分类器上平均准确率提升10.03%,比正向模式方法高24.75%。
- 适合处理复杂或不均衡的病毒基因组数据,提升可解释性。
病毒是地球上最丰富的生物实体,在微生物生态系统中扮演关键角色,同时作为重要的人类病原体,与人类发病率和死亡率密切相关。因此,从病毒基因组序列中准确识别病毒序列至关重要。然而,现有基于基因组的分类模型大多依赖组成或频率的子序列特征,常因可解释性差、在复杂或不平衡数据集上表现不佳而受限。为此,我们提出 GeneNSPCla(基于基因组负向序列模式的分类),一种新型病毒分类框架,利用负向序列模式(NSPs)从RNA病毒基因组的核苷酸序列中提取具有判别力的缺失特征。通过将这些NSPs转换为数值特征向量并整合到多个监督分类器中,GeneNSPCla 能有效捕捉病毒序列中的存在与缺失信号。此外,我们提出了一个专为基因组数据设计的负向模式挖掘算法 GONPM+,可发现更长且更具生物学意义的负向序列模式。实验结果表明,GONPM+ 在8个分类器上的平均准确率相比原始负向模式挖掘算法提升10.03%,相比正向模式挖掘算法提升24.75%。这些发现凸显了引入缺失序列信息的有效性,为病毒基因组分析与分类提供了新的互补视角。
原文摘要 · Abstract (English)
Viruses represent the most abundant biological entities on Earth and play a pivotal role in microbial ecosystems, yet, as prominent human pathogens, they are closely linked to human morbidity and mortality. Accurate identification of viral sequences from viral genome sequences is therefore essential, but existing genome-based classification models that largely relying on composition- or frequency-based subsequence features often suffer from limited interpretability and reduced accuracy, particularly on complex or imbalanced datasets. To address these limitations, we propose GeneNSPCla (Genomic Negative Sequential Pattern-based Classification), a novel viral classification framework based on Negative Sequential Patterns (NSPs) that extracts discriminative absence-based features from nucleotide sequences of RNA viral genomes. By transforming these NSPs into numerical feature vectors and integrating them into multiple supervised classifiers, GeneNSPCla effectively captures both presence and absence signals in viral sequences. Furthermore, we propose a negative pattern mining algorithm adapted for processing genomic data: GONPM+, which can discover longer and more biologically meaningful negative sequential patterns. The experimental results demonstrate that the average accuracy of GONPM+ in 8 classifiers has improved by 10.03% compared to the original negative pattern mining algorithm and by 24.75% compared to the positive pattern mining algorithm. These findings highlight the effectiveness of incorporating absence-based sequential information, providing a new and complementary perspective for viral genome analysis and classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。