arXiv:2604.11852q-bio.QMcs.AI2026-04

仅靠蛋白质序列难区分帕金森病,现有方法效果有限。

Limitations of Sequence-Based Protein Representations for Parkinson's Disease Classification: A Leakage-Free Benchmark

论文配图:Limitations of Sequence-Based Protein Representations for Parkinson's Disease Classification: A Leakage-Free Benchmark
图 1 · 摘自论文原文
  • 用纯序列数据构建多种表示法,严格避免信息泄露
  • 最佳模型F1仅0.704,多数方法在0.60-0.70间波动
  • 序列本身结构与疾病标签无关,需引入新特征

帕金森病的分子生物标志物识别因多因素特性仍具挑战。尽管蛋白质序列是基础且广泛可用的信息源,其独立用于复杂疾病分类的判别能力尚不明确。本文在嵌套分层交叉验证框架下,对仅基于蛋白质一级序列的多种表示方法(氨基酸组成、k-mers、理化描述符、混合表示及蛋白语言模型嵌入)进行了受控且无泄漏的评估。最优配置(ProtBERT + MLP)取得F1-score为0.704 ± 0.028,ROC-AUC为0.748 ± 0.047,表明判别性能仅达中等水平。经典方法如k-mers虽可达约0.667的F1值,但存在严重偏差(召回率接近0.98,精确率约0.50),倾向过度预测阳性样本。各类方法表现差异极小(F1介于0.60至0.70之间),无监督分析未发现与类别标签一致的内在结构,统计检验(Friedman检验,p=0.1749)也未显示模型间显著差异。结果表明类间高度重叠,一级序列信息本身对帕金森病分类提供有限判别力。本研究建立了可复现基线,实证支持需引入结构、功能或互作等更丰富生物特征以实现稳健建模。

原文摘要 · Abstract (English)

The identification of reliable molecular biomarkers for Parkinson's disease remains challenging due to its multifactorial nature. Although protein sequences constitute a fundamental and widely available source of biological information, their standalone discriminative capacity for complex disease classification remains unclear. In this work, we present a controlled and leakage-free evaluation of multiple representations derived exclusively from protein primary sequences, including amino acid composition, k-mers, physicochemical descriptors, hybrid representations, and embeddings from protein language models, all assessed under a nested stratified cross-validation framework to ensure unbiased performance estimation. The best-performing configuration (ProtBERT + MLP) achieves an F1-score of 0.704 +/- 0.028 and ROC-AUC of 0.748 +/- 0.047, indicating only moderate discriminative performance. Classical representations such as k-mers reach comparable F1 values (up to approximately 0.667), but exhibit highly imbalanced behavior, with recall close to 0.98 and precision around 0.50, reflecting a strong bias toward positive predictions. Across representations, performance differences remain within a narrow range (F1 between 0.60 and 0.70), while unsupervised analyses reveal no intrinsic structure aligned with class labels, and statistical testing (Friedman test, p = 0.1749) does not indicate significant differences across models. These results demonstrate substantial overlap between classes and indicate that primary sequence information alone provides limited discriminative power for Parkinson's disease classification. This work establishes a reproducible baseline and provides empirical evidence that more informative biological features, such as structural, functional, or interaction-based descriptors, are required for robust disease modeling.

蛋白质序列帕金森病生物标志物分类性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。