arXiv:2607.27143cs.LGcs.AI2026-07

解决高风险决策中少数类误判成本高的问题,让系统更懂何时该求助人类。

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

论文配图:Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
图 1 · 摘自论文原文
  • 用分层校准方法提升罕见事件的置信度覆盖,避免漏判
  • 在15个真实数据集上,少数类覆盖率平均提升61.7个百分点
  • 给出人机协同的最优拒答阈值,适合医疗、金融等高风险场景

信用评分、欺诈检测、医疗和工业安全等高风险决策系统需在严重类别不平衡和非对称错误成本下实现可靠的不确定性量化。标准边际共形预测(CP)虽能保证整体覆盖,但对代价高昂的少数类覆盖严重不足,某些数据集上仅达0.5%。本文通过涵盖15个真实世界不平衡表格数据集、7种分类模型、3种概率校准方法和10个随机种子的全面基准测试(共3150次实验),对比了边际CP、类条件(Mondrian)CP及成本控制拒答机制。结果表明,Mondrian CP恢复了有效的少数类覆盖,相比边际CP平均提升61.7个百分点(p < 1e-80)。进一步结合成本控制拒答,显著降低预期决策成本,优于标准决策边界、置信度拒答器与风险控制拒答器。同时量化了不同数据集的盈亏平衡阈值,为实际部署分布无关、成本感知的不确定性量化提供实用指导。

原文摘要 · Abstract (English)

High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To characterize and address this limitation, we conduct a comprehensive benchmark comparing marginal CP, class-conditional (Mondrian) CP, and cost-controlled abstention mechanisms across 15 real-world imbalanced tabular datasets, 7 classification models, 3 probability calibration techniques, and 10 random seeds, resulting in 3,150 experimental runs. Our results show that Mondrian CP restores valid minority-class coverage, achieving an average minority-coverage improvement of 61.7 percentage points over marginal CP (p < 1e-80). Furthermore, combining Mondrian CP with cost-controlled abstention significantly reduces expected decision cost compared with standard decision boundaries, confidence-based rejectors, and risk-controlled rejectors under realistic human review budgets. We further quantify dataset-specific break-even thresholds at which deferring ambiguous instances to human experts becomes cost-effective. These findings provide practical guidance for deploying distribution-free, cost-aware uncertainty quantification in high-stakes decision support systems.

共形预测高风险决策人机协同不平衡数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。