arXiv:2410.17293q-bio.QMcs.LG2024-10被引 1

用融合注意力的CNN-BiLSTM模型提升蛋白质家族分类精度

A Fusion-Driven Approach of Attention-Based CNN-BiLSTM for Protein Family Classification -- ProFamNet

  • 结合CNN提取空间特征、BiLSTM捕捉长程依赖、注意力聚焦关键区域
  • 参数量仅172万,比顶尖模型小90%,在27万样本上达98.3%准确率
  • 适合需要轻量化高精度蛋白质功能预测的研究者使用

先进的自动化AI技术使蛋白质序列分类及功能识别成为可能。传统方法多依赖序列中N-Gram特征,忽视了关键基序信息及其与邻近氨基酸的相互作用。近年来,卷积神经网络已被用于氨基酸和基序数据,即使在少量已知蛋白数据下也取得性能提升。本文提出ProFamNet模型,融合一维CNN、双向LSTM与注意力机制,结合空间特征提取、长期依赖建模与上下文感知表示能力。该模型仅需450,953个参数,模型大小为1.72 MB,显著小于当前最优模型(4,578,911参数,17.47 MB)。在271,160个实例上,仅用25个训练轮次即达到98.30%的F1分数,优于对比模型(97.67% F1,55,077样本,30轮训练)。

原文摘要 · Abstract (English)

Advanced automated AI techniques allow us to classify protein sequences and discern their biological families and functions. Conventional approaches for classifying these protein families often focus on extracting N-Gram features from the sequences while overlooking crucial motif information and the interplay between motifs and neighboring amino acids. Recently, convolutional neural networks have been applied to amino acid and motif data, even with a limited dataset of well-characterized proteins, resulting in improved performance. This study presents a model for classifying protein families using the fusion of 1D-CNN, BiLSTM, and an attention mechanism, which combines spatial feature extraction, long-term dependencies, and context-aware representations. The proposed model (ProFamNet) achieved superior model efficiency with 450,953 parameters and a compact size of 1.72 MB, outperforming the state-of-the-art model with 4,578,911 parameters and a size of 17.47 MB. Further, we achieved a higher F1 score (98.30% vs. 97.67%) with more instances (271,160 vs. 55,077) in fewer training epochs (25 vs. 30).

蛋白质分类深度学习生物信息学轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。