arXiv:2410.12655cs.LG2024-10被引 1

用加权位置评分核提升蛋白质序列分类准确率

Position Specific Scoring Is All You Need? Revisiting Protein Sequence Classification Tasks

  • 将位置评分与字符串核结合,构建新型核函数
  • 在多个数据集上分类准确率提升最高达45.1%
  • 适合蛋白功能预测与药物研发相关研究者

理解蛋白质的结构与功能特征对药物发现、政策制定等至关重要。位置特定评分(PSS)是分析氨基酸如何构成这些特征的重要方法。尽管字符串核在自然语言处理中至关重要,但其能否从蛋白质序列中提取生物学意义仍不明确,尽管它在一般序列分析任务中表现良好。本文提出加权位置评分核矩阵(W-PSSKM),将蛋白质序列的PSS表示(编码每个氨基酸的频率信息)与字符串核概念结合,形成一种新核函数。该方法在多项蛋白质序列分类任务中显著优于现有基线和最先进方法,分类准确率最高提升45.1%。通过广泛实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Understanding the structural and functional characteristics of proteins are crucial for developing preventative and curative strategies that impact fields from drug discovery to policy development. An important and popular technique for examining how amino acids make up these characteristics of the protein sequences with position-specific scoring (PSS). While the string kernel is crucial in natural language processing (NLP), it is unclear if string kernels can extract biologically meaningful information from protein sequences, despite the fact that they have been shown to be effective in the general sequence analysis tasks. In this work, we propose a weighted PSS kernel matrix (or W-PSSKM), that combines a PSS representation of protein sequences, which encodes the frequency information of each amino acid in a sequence, with the notion of the string kernel. This results in a novel kernel function that outperforms many other approaches for protein sequence classification. We perform extensive experimentation to evaluate the proposed method. Our findings demonstrate that the W-PSSKM significantly outperforms existing baselines and state-of-the-art methods and achieves up to 45.1\% improvement in classification accuracy.

蛋白质分类核方法序列分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。