arXiv:2605.06479stat.MLcs.LG2026-05

在不大幅改变原有决策的前提下,精准控制风险并提升政策可靠性。

Risk-Controlled Post-Processing of Decision Policies

论文配图:Risk-Controlled Post-Processing of Decision Policies
图 1 · 摘自论文原文
  • 设计阈值策略,在关键场景下切换至更安全的备用决策
  • 理论证明误差随样本量增长呈对数递减,保障稳定性能
  • 适合医疗、AI路由等高风险需谨慎调整的决策系统

预测模型常通过现有决策策略部署,利益相关方不愿轻易更改。本文研究风险可控的后处理:给定一个确定性基线策略,选择新策略以最大化与基线的一致性,同时满足用户指定损失的随机约束。在总体层面,最优策略具有阈值结构:仅在切换至理想备选策略能显著降低条件违规风险时才偏离基线。在有限样本情况下,利用校准数据和拟合的备选策略与评分,提出一种后处理算法来选取阈值。借助算法稳定性与随机过程工具,在独立同分布条件下,后处理策略的期望超额风险为 $O(\log n/n)$。当存在精确安全的备选策略时,算法可在可交换性假设下实现精确的期望风险控制,并提供高概率近优保证。在新冠肺部影像诊断、大模型路由及合成多分类任务上的实验表明,目标化后处理能在满足或接近风险预算的同时,比盲目的随机混合保留更多与基线的一致性。

原文摘要 · Abstract (English)

Predictive models are often deployed through existing decision policies that stakeholders are reluctant to change unless a risk constraint requires intervention. We study risk-controlled post-processing: given a deterministic baseline policy, choose a new policy that maximizes agreement with the baseline subject to a chance constraint on a user-specified loss. At the population level, we show that the optimal policy has a threshold structure: it follows the baseline except on contexts where switching to the oracle fallback policy yields a large reduction in conditional violation risk. At the finite-sample level, given a fitted fallback policy and score, we develop a post-processing algorithm that uses calibration data to select a threshold. Leveraging tools from algorithmic stability and stochastic processes, we show that under regularity conditions, in the i.i.d. setting, the expected excess risk of the post-processed policy is $O(\log n/n)$. In the special case when an exact-safe fallback policy is available, the algorithm achieves precise expected risk control under exchangeability. In this setting, we also give high-probability near-optimality guarantees on the post-processed policy. Experiments on a COVID-19 radiograph diagnosis task, an LLM routing problem, and a synthetic multiclass decision task show that targeted post-processing can meet or nearly meet risk budgets while preserving substantially more agreement with the baseline than score-blind random mixing.

决策优化风险控制后处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。