arXiv:2603.09947cs.AI2026-03

提出信心门控的适用条件,帮系统判断何时该放弃决策。

The Confidence Gate Theorem: When Should Ranked Decision Systems Abstain?

  • 用结构与上下文不确定性区分信心失效场景
  • 结构不确定性下信心门控可提升决策质量,上下文漂移则会失效
  • 推荐、电商、医疗三领域验证,提醒选对信心信号

排名决策系统(如推荐、广告拍卖、临床分诊)需决定何时干预或放弃输出。本文研究基于信心的弃权机制何时能单调提升决策质量,何时会失败。理论条件清晰:排名对齐且无反转区。关键发现是,这取决于结构性不确定性(如冷启动)与上下文不确定性(如时间漂移)的差异。实证在三个领域验证:协同过滤(MovieLens,3种分布偏移)、电商意图识别(RetailRocket、Criteo、Yoochoose)、临床路径分诊(MIMIC-IV)。结构性不确定性下,信心门控近乎单调增益;而基于观察次数的信心信号在上下文漂移下表现堪比随机弃权(MovieLens时序分裂中出现同等数量的单调性破坏)。上下文感知替代方案(集成分歧、新近特征)显著缩小差距(违规从3降至1–2),但未完全恢复单调性,表明上下文不确定性带来本质挑战。基于残差定义的异常标签在分布偏移下性能大幅下降(AUC从0.71降至0.61–0.62),否定异常驱动干预的常见做法。结果提供实用部署诊断:部署前在保留数据上检验条件C1和C2,确保信心信号与主导不确定性类型匹配。

原文摘要 · Abstract (English)

Ranked decision systems -- recommenders, ad auctions, clinical triage queues -- must decide when to intervene in ranked outputs and when to abstain. We study when confidence-based abstention monotonically improves decision quality, and when it fails. The formal conditions are simple: rank-alignment and no inversion zones. The substantive contribution is identifying why these conditions hold or fail: the distinction between structural uncertainty (missing data, e.g., cold-start) and contextual uncertainty (missing context, e.g., temporal drift). Empirically, we validate this distinction across three domains: collaborative filtering (MovieLens, 3 distribution shifts), e-commerce intent detection (RetailRocket, Criteo, Yoochoose), and clinical pathway triage (MIMIC-IV). Structural uncertainty produces near-monotonic abstention gains in all domains; structurally grounded confidence signals (observation counts) fail under contextual drift, producing as many monotonicity violations as random abstention on our MovieLens temporal split. Context-aware alternatives -- ensemble disagreement and recency features -- substantially narrow the gap (reducing violations from 3 to 1--2) but do not fully restore monotonicity, suggesting that contextual uncertainty poses qualitatively different challenges. Exception labels defined from residuals degrade substantially under distribution shift (AUC drops from 0.71 to 0.61--0.62 across three splits), providing a clean negative result against the common practice of exception-based intervention. The results provide a practical deployment diagnostic: check C1 and C2 on held-out data before deploying a confidence gate, and match the confidence signal to the dominant uncertainty type.

决策系统信心门控不确定性推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。