提出新评估方法,更好衡量不确定信息增强系统在高风险决策中的表现
ECUAS$_n$: A family of metrics for principled evaluation of uncertainty-augmented systems
- 基于合理评分规则设计统一评估指标,融合预测与不确定性
- 参数n可调节错误预测与不确定性偏差的权衡,适配不同应用场景
- 在分类与生成任务上验证有效,支持手动标注数据集测试
在高风险自动化决策中,获取预测不确定性对用户(人类或下游系统)根据具体成本权衡接受或拒绝预测至关重要。当前文献对不确定性增强(UA)系统——即输出预测和不确定性评分的系统——的评估方式各异,常采用独立指标评价预测与不确定性,设定固定拒收成本,或积分覆盖-风险曲线。我们认为这些方法无法全面评估UA系统在不确定性下的决策性能,因此提出一类新指标ECUAS$_n$,其形式为针对具体任务的严格评分规则。参数$n$可根据应用需求调节错误预测与不精确不确定性之间的权衡。我们通过在多样化的分类与生成数据集上的实验,包括TriviaQA的手动标注子集,从理论和实证两方面展示了ECUAS$_n$的优势。
原文摘要 · Abstract (English)
In high-stakes automated decision-making, access to predictive uncertainty is essential for enabling users -- human or downstream systems -- to accept or reject predictions based on application-specific cost trade-offs. Such uncertainty-augmented (UA) systems -- i.e., systems that output both predictions and uncertainty scores -- are currently being assessed in the literature in a variety of ways, using separate metrics to evaluate the predictions and the uncertainty scores, setting a cost function with a fixed rejection cost or integrating over a coverage-risk curve. We argue that these evaluation approaches are inadequate for assessing overall performance of the UA system for decision making under uncertainty and propose a novel family of metrics, ECUAS$_n$, formulated as proper scoring rules for the task of interest. The parameter $n$ controls the trade-off between the cost of incorrect predictions and imperfect uncertainties depending on the needs of the use-case. We demonstrate the advantages of the ECUAS$_n$ metrics both theoretically and empirically, through experiments on diverse classification and generation datasets, including a manually annotated subset of TriviaQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。