用极端值理论评估高风险场景下的模型灾难性失败概率。
New Statistical Framework for Extreme Error Probability in High-Stakes Domains for Reliable Machine Learning
- 基于极值理论构建新统计框架,量化罕见但严重的预测错误。
- 在真实与合成数据上验证,可稳健估计灾难性故障发生概率。
- 适合需要可靠安全性的医疗、自动驾驶等高风险领域使用。
机器学习在高风险领域至关重要,但传统验证方法依赖均方误差(MSE)或平均绝对误差(MAE)等平均指标,无法衡量极端错误。最坏情况下的预测失败可能带来严重后果,而现有框架缺乏对其实现概率的统计基础。本文提出一种基于极值理论(EVT)的新统计框架,为评估最坏情况失败提供严格方法。在合成与真实数据集上的应用表明,该方法能有效克服标准交叉验证的根本局限,实现对灾难性故障概率的稳健估计。本工作确立了极值理论作为评估模型可靠性的重要工具,确保在不确定性量化至关重要的新技术部署中实现更安全的人工智能应用。
原文摘要 · Abstract (English)
Machine learning is vital in high-stakes domains, yet conventional validation methods rely on averaging metrics like mean squared error (MSE) or mean absolute error (MAE), which fail to quantify extreme errors. Worst-case prediction failures can have substantial consequences, but current frameworks lack statistical foundations for assessing their probability. In this work a new statistical framework, based on Extreme Value Theory (EVT), is presented that provides a rigorous approach to estimating worst-case failures. Applying EVT to synthetic and real-world datasets, this method is shown to enable robust estimation of catastrophic failure probabilities, overcoming the fundamental limitations of standard cross-validation. This work establishes EVT as a fundamental tool for assessing model reliability, ensuring safer AI deployment in new technologies where uncertainty quantification is central to decision-making or scientific analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。