弱校准预测下如何稳健决策,找到最坏情况下的最优行动策略。
Robust Decision Making with Partially Calibrated Forecasts
- 基于对偶分析推导出最小最大最优决策规则
- 即使校准不足,仍可高效计算最优策略
- 适合需在不确定预测中做稳健决策的研究者
校准是可信机器学习中的基础目标,因其强决策论意义:无论分布或效用函数如何,充分校准的预测下,最优策略即“信任预测并据此行动”。但充分校准仅适用于低维问题。高维问题(如多分类)中常用弱校准形式,缺乏此类决策优势。本文研究保守决策者如何将具有弱校准保证的预测映射为行动,在分布一致的前提下最大化最坏情况下的期望效用。通过对偶论证刻画了最小最大最优决策规则,发现“信任预测”在决策校准(decision calibration)及更强条件下的最小最大意义下依然成立——该条件远弱于充分校准且更易实现。对于未达决策校准的校准保证,最优决策规则仍可高效计算,并提供一种适用于任意平方误差优化回归模型的自然方法的实证评估。
原文摘要 · Abstract (English)
Calibration has emerged as a foundational goal in ``trustworthy machine learning'', in part because of its strong decision theoretic semantics. Independent of the underlying distribution, and independent of the decision maker's utility function, calibration promises that amongst all policies mapping predictions to actions, the uniformly best policy is the one that ``trusts the predictions'' and acts as if they were correct. But this is true only of \emph{fully calibrated} forecasts, which are tractable to guarantee only for very low dimensional prediction problems. For higher dimensional prediction problems (e.g. when outcomes are multiclass), weaker forms of calibration have been studied that lack these decision theoretic properties. In this paper we study how a conservative decision maker should map predictions endowed with these weaker (``partial'') calibration guarantees to actions, in a way that is robust in a minimax sense: i.e. to maximize their expected utility in the worst case over distributions consistent with the calibration guarantees. We characterize their minimax optimal decision rule via a duality argument, and show that surprisingly, ``trusting the predictions and acting accordingly'' is recovered in this minimax sense by \emph{decision calibration} (and any strictly stronger notion of calibration), a substantially weaker and more tractable condition than full calibration. For calibration guarantees that fall short of decision calibration, the minimax optimal decision rule is still efficiently computable, and we provide an empirical evaluation of a natural one that applies to any regression model solved to optimize squared error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。