提出风险敏感的分类评估指标,降低高自信错误带来的风险。
Fragility-aware Classification for Understanding Risk and Improving Generalization
- 引入脆弱性指数FI,衡量高自信错误的尾部风险
- 实验证明FI模型显著减少高自信误判,成本更低
- 适用于医疗诊断等高风险决策场景
分类模型在医疗诊断、推荐系统和风险评估等数据驱动决策中起核心作用。传统指标如准确率和AUC关注整体误差率,却忽视了错误预测的置信度,即高自信误判的风险。这一缺陷在安全关键和成本敏感场景中尤为严重,因过度自信的错误可能导致严重后果。为此,我们提出脆弱性指数(Fragility Index, FI),一种从风险规避角度评估分类器的新指标,捕捉高自信误判的尾部风险。我们在鲁棒满意(RS)框架下构建FI,确保分布不确定性下的稳健性。基于此,我们设计可训练的优化框架,通过代理损失直接优化FI,并证明该框架下模型具备可证明的FI边界。我们还为交叉熵、铰链型及Lipschitz损失等广泛损失函数推导出精确重构形式,并扩展至深度神经网络。真实世界医疗诊断任务的实验表明,FI弥补了现有指标的不足,揭示了误差尾部风险,提升了决策质量。基于FI的模型在保持竞争力准确率和AUC的同时,持续降低高自信误判及其运营成本,为高风险应用中的鲁棒性和可靠性提供实用工具。
原文摘要 · Abstract (English)
Classification models play a central role in data-driven decision-making applications such as medical diagnosis, recommendation systems, and risk assessment. Traditional performance metrics, such as accuracy and AUC, focus on overall error rates but fail to account for the confidence of incorrect predictions, i.e., the risk of confident misjudgments. This limitation is particularly consequential in safety-critical and cost-sensitive settings, where overconfident errors can lead to severe outcomes. To address this issue, we propose the Fragility Index (FI), a novel performance metric that evaluates classifiers from a risk-averse perspective by capturing the tail risk of confident misjudgments. We formulate FI within a robust satisficing (RS) framework to ensure robustness under distributional uncertainty. Building on this, we develop a tractable training framework that directly targets FI via a surrogate loss, and show that models trained under this framework admit provable bounds on FI. We further derive exact reformulations for a broad class of loss functions, including cross-entropy, hinge-type, and Lipschitz losses, and extend the approach to deep neural networks. Empirical results on real-world medical diagnosis tasks demonstrate that FI complements existing metrics by revealing error tail risk and improving decision quality. FI-based models achieve competitive accuracy and AUC while consistently reducing confident misjudgments and associated operational costs, offering a practical tool for improving robustness and reliability in risk-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。