arXiv:2411.12853cs.LGq-bio.BM2024-11

将二级结构信息融入空间关系模型,提升蛋白质分类精度。

Integrating Secondary Structures Information into Triangular Spatial Relationships (TSR) for Advanced Protein Classification

  • 在三角空间关系中引入螺旋、链、无规卷曲的18种组合
  • 在两个数据集上分别实现98.3%和99.5%的分类准确率
  • 特别适合低准确率数据集,对高基线数据增益有限

蛋白质结构是解读生物功能的关键。传统结构比对方法常忽略更细粒度的相似性,而先进方法如三角空间关系(TSR)能实现更精细区分。但经典TSR未融合二级结构信息,影响对折叠模式的理解。为此,我们提出SSE-TSR方法,将二级结构元件(SSEs)融入基于TSR的蛋白质表示,考虑18种螺旋、链、无规卷曲的排列组合。实验使用两个大规模数据集(9.2K和7.8K样本),结合神经网络进行分类。结果显示,引入SSE后,数据集1的准确率从96.0%提升至98.3%,数据集2从99.4%微增至99.5%。表明该方法在初始性能较低的数据集中效果显著,在高基线数据中仍有小幅提升。因此,SSE-TSR是一种有效提升蛋白质分类与功能理解的生物信息学工具。

原文摘要 · Abstract (English)

Protein structures represent the key to deciphering biological functions. The more detailed form of similarity among these proteins is sometimes overlooked by the conventional structural comparison methods. In contrast, further advanced methods, such as Triangular Spatial Relationship (TSR), have been demonstrated to make finer differentiations. Still, the classical implementation of TSR does not provide for the integration of secondary structure information, which is important for a more detailed understanding of the folding pattern of a protein. To overcome these limitations, we developed the SSE-TSR approach. The proposed method integrates secondary structure elements (SSEs) into TSR-based protein representations. This allows an enriched representation of protein structures by considering 18 different combinations of helix, strand, and coil arrangements. Our results show that using SSEs improves the accuracy and reliability of protein classification to varying degrees. We worked with two large protein datasets of 9.2K and 7.8K samples, respectively. We applied the SSE-TSR approach and used a neural network model for classification. Interestingly, introducing SSEs improved performance statistics for Dataset 1, with accuracy moving from 96.0% to 98.3%. For Dataset 2, where the performance statistics were already good, further small improvements were found with the introduction of SSE, giving an accuracy of 99.5% compared to 99.4%. These results show that SSE integration can dramatically improve TSR key discrimination, with significant benefits in datasets with low initial accuracies and only incremental gains in those with high baseline performance. Thus, SSE-TSR is a powerful bioinformatics tool that improves protein classification and understanding of protein function and interaction.

蛋白质分类二级结构空间关系生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。