提出新方法量化预测不确定性,识别高风险特征区间。
Model uncertainty quantification using feature confidence sets for outcome excursions
- 用特征空间的内外置信集捕捉结果超阈值的区域
- 理论保证在有限样本下仍能覆盖真实高风险区域
- 适合医疗、金融等高风险场景的决策支持
在医疗、金融、自动驾驶等高风险应用中,预测不确定性量化对风险管理至关重要。传统方法如置信区间和预测区间仅针对期望结果或实际结果提供概率覆盖。本文提出一种模型无关的新框架,通过构建结果超出特定阈值的特征空间置信集,实现对连续与二分类结果的不确定性量化。该方法生成数据依赖的内、外置信集,旨在包含真实使期望或实际结果超过阈值的特征子集。研究建立了渐近及有限样本下的理论保证,证明置信集包含真实区域的概率。通过仿真和真实数据验证,展示了在房价预测与败血症诊断时间预测中的有效性。该方法为各类预测模型提供了统一的不确定性量化手段。
原文摘要 · Abstract (English)
When implementing prediction models for high-stakes real-world applications such as medicine, finance, and autonomous systems, quantifying prediction uncertainty is critical for effective risk management. Traditional approaches to uncertainty quantification, such as confidence and prediction intervals, provide probability coverage guarantees for the expected outcomes $f(\boldsymbol{x})$ or the realized outcomes $f(\boldsymbol{x})+ε$. Instead, this paper introduces a novel, model-agnostic framework for quantifying uncertainty in continuous and binary outcomes using confidence sets for outcome excursions, where the goal is to identify a subset of the feature space where the expected or realized outcome exceeds a specific value. The proposed method constructs data-dependent inner and outer confidence sets that aim to contain the true feature subset for which the expected or realized outcomes of these features exceed a specified threshold. We establish theoretical guarantees for the probability that these confidence sets contain the true feature subset, both asymptotically and for finite sample sizes. The framework is validated through simulations and applied to real-world datasets, demonstrating its utility in contexts such as housing price prediction and time to sepsis diagnosis in healthcare. This approach provides a unified method for uncertainty quantification that is broadly applicable across various continuous and binary prediction models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。