arXiv:2512.00376q-bio.QMcs.AI2025-12被引 2

通过分析Transformer中间层提升激酶功能预测准确率

Layer Probing Improves Kinase Functional Prediction with Protein Language Models

  • 从ESM-2模型33层中筛选中到深层的特征进行预测
  • 中深层层使无监督评估指标提升32%,有监督准确率达75.7%
  • 适合关注蛋白质语言模型深层信号的研究者

蛋白质语言模型(PLMs)已革新基于序列的蛋白分析,但多数应用仅依赖最终层嵌入,可能忽略早期层中的生物学信息。我们系统评估了ESM-2模型全部33层在激酶功能预测中的表现,采用无监督聚类和有监督分类两种方式。结果表明,中至深层(第20-33层)相比最终层,在无监督调整兰德指数上提升32%,并使同源感知的有监督准确率达到75.7%。通过领域级特征提取、校准概率估计及可复现的基准测试流程,进一步提升了可靠性。研究证明,Transformer深度中包含功能各异的生物信号,而合理选择层数能显著提升激酶功能预测性能。

原文摘要 · Abstract (English)

Protein language models (PLMs) have transformed sequence-based protein analysis, yet most applications rely only on final-layer embeddings, which may overlook biologically meaningful information encoded in earlier layers. We systematically evaluate all 33 layers of ESM-2 for kinase functional prediction using both unsupervised clustering and supervised classification. We show that mid-to-late transformer layers (layers 20-33) outperform the final layer by 32 percent in unsupervised Adjusted Rand Index and improve homology-aware supervised accuracy to 75.7 percent. Domain-level extraction, calibrated probability estimates, and a reproducible benchmarking pipeline further strengthen reliability. Our results demonstrate that transformer depth contains functionally distinct biological signals and that principled layer selection significantly improves kinase function prediction.

蛋白质语言模型激酶预测Transformer层分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。