arXiv:2507.20426cs.LGcs.AI2025-07被引 1

用蛋白嵌入+轻量胶囊网络,精准预测DNA结合蛋白

ResCap-DBP: A Lightweight Residual-Capsule Network for Accurate DNA-Binding Protein Prediction Using Global ProteinBERT Embeddings

  • 残差胶囊结构融合全局蛋白嵌入,捕获序列深层特征
  • 在4个数据集上AUC最高达98.0%,小样本也表现稳定
  • 适合基因功能研究与疾病机制探索者使用

DNA结合蛋白(DBPs)在基因调控和细胞过程中起关键作用,其准确识别对理解生物功能和疾病机制至关重要。实验方法耗时且成本高,亟需高效计算预测技术。本文提出ResCap-DBP框架,结合基于残差学习的编码器与一维胶囊网络(1D-CapsNet),直接从原始蛋白序列预测DBPs。残差块中引入空洞卷积缓解梯度消失,胶囊层通过动态路由捕捉特征空间中的层次与空间关系。对比ProteinBERT的全局/局部嵌入与传统one-hot编码的消融实验显示,ProteinBERT嵌入在大数据集上显著优于其他表示;尽管one-hot在小数据集(如PDB186)略有优势,但难以扩展。在四个公开基准数据集上的评估表明,本模型持续超越当前最优方法:在PDB14189和PDB1075上分别达到98.0%和89.5%的AUC;在独立测试集PDB2272和PDB186上分别取得83.2%和83.3%的AUC,同时在大规模数据集(如PDB20000)上保持良好性能。模型在各数据集上均保持良好的敏感性与特异性平衡,验证了将全局蛋白表示与先进深度架构结合,在多样基因组背景下实现可靠、可扩展的DBP预测的有效性。

原文摘要 · Abstract (English)

DNA-binding proteins (DBPs) are integral to gene regulation and cellular processes, making their accurate identification essential for understanding biological functions and disease mechanisms. Experimental methods for DBP identification are time-consuming and costly, driving the need for efficient computational prediction techniques. In this study, we propose a novel deep learning framework, ResCap-DBP, that combines a residual learning-based encoder with a one-dimensional Capsule Network (1D-CapsNet) to predict DBPs directly from raw protein sequences. Our architecture incorporates dilated convolutions within residual blocks to mitigate vanishing gradient issues and extract rich sequence features, while capsule layers with dynamic routing capture hierarchical and spatial relationships within the learned feature space. We conducted comprehensive ablation studies comparing global and local embeddings from ProteinBERT and conventional one-hot encoding. Results show that ProteinBERT embeddings substantially outperform other representations on large datasets. Although one-hot encoding showed marginal advantages on smaller datasets, such as PDB186, it struggled to scale effectively. Extensive evaluations on four pairs of publicly available benchmark datasets demonstrate that our model consistently outperforms current state-of-the-art methods. It achieved AUC scores of 98.0% and 89.5% on PDB14189andPDB1075, respectively. On independent test sets PDB2272 and PDB186, the model attained top AUCs of 83.2% and 83.3%, while maintaining competitive performance on larger datasets such as PDB20000. Notably, the model maintains a well balanced sensitivity and specificity across datasets. These results demonstrate the efficacy and generalizability of integrating global protein representations with advanced deep learning architectures for reliable and scalable DBP prediction in diverse genomic contexts.

蛋白质预测深度学习胶囊网络基因调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。