为预测模型提供无需假设分布的可靠性评分,识别高风险预测。
LOCUS: A Distribution-Free Loss-Quantile Score for Risk-Aware Predictions
- 用输入条件下的损失分布建模,生成可跨样本比较的可靠性分数。
- 在13个回归任务中显著降低大误差事件频率,优于传统方法。
- 适合需要可控高风险预警的工业部署场景。
现代机器学习模型虽平均精度高,但仍可能产生导致部署成本飙升的错误。本文提出LOCUS,一种无需分布假设的封装工具,为固定预测函数生成每个输入的损失尺度可靠性评分。不同于对标签不确定性的建模,LOCUS通过任意能输出给定输入下损失预测分布的引擎,建模预测函数的实际损失。经简单分段校准后,该评分具备分布无关性、可解释性,且可直接解读为损失上限。评分本身可用于排序,也可设定阈值实现透明的异常预警,分布无关地控制大损失事件。在13个回归基准测试中,LOCUS展现出有效风险排序能力,并显著降低大损失频率,优于标准启发式方法。
原文摘要 · Abstract (English)
Modern machine learning models can be accurate on average yet still make mistakes that dominate deployment cost. We introduce Locus, a distribution-free wrapper that produces a per-input loss-scale reliability score for a fixed prediction function. Rather than quantifying uncertainty about the label, Locus models the realized loss of the prediction function using any engine that outputs a predictive distribution for the loss given an input. A simple split-calibration step turns this function into a distribution-free interpretable score that is comparable across inputs and can be read as an upper loss level. The score is useful on its own for ranking, and it can optionally be thresholded to obtain a transparent flagging rule with distribution-free control of large-loss events. Experiments across 13 regression benchmarks show that Locus yields effective risk ranking and reduces large-loss frequency compared to standard heuristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。