arXiv:2609.07897cs.LGcs.AI2026-09

默认阈值导致酶分类预测严重失准,真实效果远低于表面准确率。

The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]

  • 用实证方法诊断多标签酶分类中固定阈值的缺陷
  • 平均准确率77.16%但少数类召回率低至0%,EC6完全失效
  • 适合生物信息学和药物发现中的模型可靠性研究者

自动化预测酶分类(EC)数在功能注释与计算药物发现中至关重要。然而,标准多标签机器学习流程常依赖默认决策阈值(t=0.50),假设各目标类别先验均衡。本研究对N=14,096个标注化合物在六类主要EC类别(EC1-EC6)下的未校准固定决策边界进行了系统性实证诊断。结果显示显著的准确率悖论:尽管平均准确率达77.16%,但宏平均F1分数仅0.3976,宏平均召回率仅为0.3872,揭示严重预测失效。多数类别出现过度敏感与过预测,少数类别召回率急剧下降,最终导致EC6决策边界彻底崩溃(召回率=0.00%),尽管其判别能力尚存(ROC-AUC=0.5857)。特征相关性分析显示拓扑指数与指纹密度度量间存在高度线性冗余。本研究证明标准点预测掩盖了生物信息学工作流中的关键错误,提出目标特异性阈值优化与事后共形校准作为必要、开源的后处理保障,以确保可靠的应用型机器学习与深度学习架构。

原文摘要 · Abstract (English)

Automated prediction of Enzyme Commission (EC) numbers plays a central role in functional annotation and computational drug discovery. However, standard multi-label machine learning pipelines frequently rely on default decision thresholds (t=0.50), assuming balanced prior distributions across target heads. In this study, we present a systematic empirical diagnostic of uncalibrated fixed decision boundaries operating under severe class imbalance across N = 14,096 annotated compounds categorized into six primary EC classes (EC1-EC6). Our results highlight a pronounced Accuracy Paradox: while the multi-label system achieves a deceivingly high mean accuracy of 77.16%, the macro F1-score (0.3976) and macro recall (0.3872) reveal severe predictive breakdown. Majority target classes suffer from hyper-sensitivity and over-prediction, whereas minority classes exhibit sharp recall decay, culminating in a total decision boundary collapse for EC6 (Recall = 0.00%) despite underlying discriminative power (ROC-AUC = 0.5857). Feature correlation analysis further reveals high linear redundancy among topological indices relative to fingerprint density metrics. Ultimately, this diagnostic study demonstrates that standard point predictions mask critical errors in bioinformatics workflows. We establish target-specific threshold optimization and post-hoc conformal calibration as essential, open-source post-processing safeguards for reliable applied machine learning and deep learning architectures.

酶分类多标签学习模型校准生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。