arXiv:2608.10893cs.CL2026-08

提出跨模型风险控制框架,实现分布偏移下精准拒答与可信覆盖率

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

  • 构建覆盖底线的认证映射,结合源域标注与目标域无标签数据
  • 在分布偏移下实现0误判拒答,1024单元审计中零违规
  • 适合需要高可靠性拒答的自动化系统部署

认证型选择性预测器在给定覆盖率下运行;操作者设定自动化底线:在分布偏移的目标流量中,至少回答β比例样本,且错误率不超过α。在有界比率分布偏移下,我们证明了‘底线认证映射’:当底线必须与选择条件下的风险α一同认证时,认证获得可行性边界与双资源复杂度图,加法可分解为:标签源域的风险项与无标签目标域的底线项。速率局部化,需满足正则边界裕度、局部范式阈值以下松弛及格点条件:上界预注册格点裕度,下界按松弛兼容。所展示的分割即为实际操作路径;若使用预言权重,还可估计源域的底线。三种模型结果:下界(Model-B),匹配预言权重上界(Model-A),以及可在预注册精确分层偏移模型下实现的上界(Model-B')。匹配是跨模型而非单模型极小极大定理,且必然如此:在全有界比率类中,任何未知权重过程在任意样本量下均无法匹配(Model-B在α=β=1/2时不一致)。干扰项必要性尚未完全确定。复杂度追踪局部接受区域泛函,而非全局有效样本量(ESS);固定ESS分离定理仍待开放;当β→0时,两条下界轴均消失,因此底线生成该映射。实验上,注册咬合族在对数-对数斜率-2.002范围内发散;1,024单元审计记录0次违规;单语料SQuAD到NewsQA可行性审计返回诚实拒答。

原文摘要 · Abstract (English)

Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a $β$-fraction of shifted target traffic with at most an $α$-fraction of answers wrong. Under bounded-ratio covariate shift we prove the Floor Certification Map: once that floor must be certified alongside the selection-conditioned risk $α$, certification acquires a feasibility frontier and a two-resource complexity map, additive up to constants: risk in labeled source, the floor in unlabeled target samples. The rates are local, needing a regular frontier margin, slack below the local-regime threshold, and lattice conditions: pre-registered with a lattice margin for the upper bounds, compatible per-slack for the lower. The displayed split is the operational route; oracle weights also allow a labeled-source floor estimate. Three model-tagged results: a lower bound (Model-B), a matching oracle-weight upper bound (Model-A), and an implementable upper bound (Model-B') valid under a pre-registered exact stratified-shift model with nuisance cost priced explicitly. The match is across these models rather than a single-model minimax theorem, and necessarily so: over the full bounded-ratio class no unknown-weight procedure matches at any sample size (Model-B is inconsistent, witnessed at $α=β=1/2$). The nuisance's necessity is only partially settled. Complexity tracks a localized accepted-region functional, not global effective sample size (ESS), on both sides, though a fixed-ESS separation theorem is left open; both lower-bound axes vanish as $β\to0$, so the floor creates the map. Empirically, the registered bite family diverges with log-log slope $-2.002$ within its pre-registered band; a 1,024-cell audit records 0 violations where the formal certificates fire; and a single-corpus SQuAD-to-NewsQA feasibility audit returns honest refusal.

风险控制分布偏移拒答机制认证学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。