arXiv:2608.08984cs.LG2026-08

解决不平衡分类中阈值依赖问题,用概率校准提升模型选择稳定性。

SoftMCC: An MCC-Brier Calibration Bridge for Threshold-Free Model Selection under Class Imbalance

论文配图:SoftMCC: An MCC-Brier Calibration Bridge for Threshold-Free Model Selection under Class Imbalance
图 1 · 摘自论文原文
  • 基于概率混淆矩阵构建软化MCC评分,避免阈值敏感性。
  • 在18个场景中稳定排名第一(均值排名2.31),相关性达0.659。
  • 适合关注校准性与模型选择一致性的研究者使用。

不平衡二分类中的模型选择常依赖马修斯相关系数(MCC),但传统阈值化使验证排名受阈值影响。本文提出SoftMCC,一种基于概率值混淆统计的后训练验证框架,结合专用校准恒等式与共享池、抗纠缠的选择协议。其核心为协方差归一化的概率-标签关联度量,对硬预测精确退化为MCC,且具有皮尔逊界。在理想校准下,其等于布里尔技能得分并保持相同排序;校准偏差时,二者差距不反映校准误差。在18组设置下进行12次重复分组测试,SoftMCC取得最佳平均排名(2.31)和最高修正肯德尔W值(0.659),弗里德曼检验显著(p=0.007);Nemenyi分析显示其显著优于AUPRC与[email protected],而14源族敏感性仅保留后者。所选模型实用性无优势。六项预设比较中三项测试MCC均值差异为负,仅F1@best通过霍尔姆校正(p=0.014),数据集级检验不显著(p=0.117)。标签置换使平均W降至0.092;温度缩放改变SoftMCC排名(均值斯皮尔曼0.851),而基于排名与阈值优化的指标保持不变。SoftMCC是一种校准敏感的MCC家族选择器,具备有界稳定性与实用性证据。

原文摘要 · Abstract (English)

Model selection for imbalanced binary classification often uses the Matthews correlation coefficient (MCC), but thresholding makes validation rankings threshold-dependent. SoftMCC is a post-training MCC validation framework on established probability-valued confusion counts, coupling an MCC-specific calibrated identity with a tie-aware, shared-pool selection protocol. Its core score is a covariance-normalized probability-label association, reduces exactly to MCC for hard predictions, and is Pearson-bounded. Under perfect population calibration it equals the Brier skill score with identical candidate ordering; outside that regime the gap does not identify calibration error. Across 18 settings with 12 duplicate-safe grouped repeats, SoftMCC attains the best stability mean rank (2.31) and highest mean tie-corrected Kendall's W (0.659), with a significant Friedman test (p=0.007); Nemenyi analysis separates it from AUPRC and [email protected], while 14-source-family sensitivity retains only the latter. Selected-model utility shows no advantage. Three of six prespecified comparisons have negative mean test-MCC differences, only F1@best survives Holm correction (p=0.014), and the dataset-level test is not significant (p=0.117). Label permutation lowers mean W to 0.092; temperature scaling shifts SoftMCC rankings (mean Spearman 0.851) whereas rank-based and threshold-optimized metrics remain invariant. SoftMCC is a calibration-sensitive MCC-family selector with bounded stability and utility evidence.

不平衡分类模型选择概率校准MCC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。