提出CCP-NN方法,高效处理分子序列分析,提升分类准确率并加速计算。
Nearest Neighbor CCP-Based Molecular Sequence Analysis
- 基于近邻搜索重构序列相关性,避免矩阵对角化,适配多种机器学习任务。
- 在分子序列分类中,准确率显著提升,计算速度比原CCP快得多。
- 适合生物信息学中大规模序列数据的预处理与分类研究者使用。
分子序列分析对于理解蛋白质互作、功能注释和疾病分类等生物过程至关重要。由于序列数量庞大且蛋白结构复杂,分析难度高,需借助降维与特征选择方法。近期提出的相关聚类与投影(CCP)方法在序列可视化方面表现有效,但计算成本高,且在序列分类中的应用尚不明确。为此,本文提出基于最近邻的CCP-NN方法,用于高效预处理分子序列数据。该方法利用序列间相关性分组相似序列,并生成代表性超序列。与传统方法不同,CCP无需矩阵对角化,可适用于多种机器学习场景。通过近邻搜索估计密度图并计算相关性。实验表明,基于CCP与CCP-NN表示的分子序列分类任务中,CCP-NN显著提升分类准确率,且计算耗时远低于原版CCP。
原文摘要 · Abstract (English)
Molecular sequence analysis is crucial for comprehending several biological processes, including protein-protein interactions, functional annotation, and disease classification. The large number of sequences and the inherently complicated nature of protein structures make it challenging to analyze such data. Finding patterns and enhancing subsequent research requires the use of dimensionality reduction and feature selection approaches. Recently, a method called Correlated Clustering and Projection (CCP) has been proposed as an effective method for biological sequencing data. The CCP technique is still costly to compute even though it is effective for sequence visualization. Furthermore, its utility for classifying molecular sequences is still uncertain. To solve these two problems, we present a Nearest Neighbor Correlated Clustering and Projection (CCP-NN)-based technique for efficiently preprocessing molecular sequence data. To group related molecular sequences and produce representative supersequences, CCP makes use of sequence-to-sequence correlations. As opposed to conventional methods, CCP doesn't rely on matrix diagonalization, therefore it can be applied to a range of machine-learning problems. We estimate the density map and compute the correlation using a nearest-neighbor search technique. We performed molecular sequence classification using CCP and CCP-NN representations to assess the efficacy of our proposed approach. Our findings show that CCP-NN considerably improves classification task accuracy as well as significantly outperforms CCP in terms of computational runtime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。