AI监督中,平滑评分无法保证诚实报告,必须用突变阈值才能避免偏差。
The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting
- 用非线性审批函数筛选类型,但会破坏报告准确性
- 任何不可检测的偏离都会导致诚实报告不再最优
- 阶梯式审批阈值可实现最优筛选,尤其对Brier得分有效
在可扩展的人工智能监管中,委托人使用严格恰当的评分规则评估代理人的报告,但代理人也通过非准确性渠道(如批准自主行动、分配份额、下游控制)获益。这种结构同样出现在经典机制设计场景中。主要结论是:委托人最优监管必然采用非线性审批函数进行类型筛选,但任何非线性审批都会使在不可检测偏离下诚实报告次优。这一不可能性对所有严格恰当的评分规则成立,并有闭式扰动公式。构造性解法存在:阶梯式审批阈值能为任意严格恰当评分规则实现第一优筛选,因为代理人的二元膨胀或不膨胀选择在类型空间形成阈值,与生成器曲率无关。针对Brier得分,类型无关的膨胀成本使得次优与第一优福利等价;我们证明该等价仅对Brier成立(对所有非Brier规则,平滑$C^1$监管下的福利差距下界为$Ω(\text{Var}(1/G'') (γ/β)^2)$)。两个实例展开框架:人工智能代理监管(核心动机场景)和市场运作(平行机制设计领域)。对人工智能对齐的启示明确:平滑评分监管无法从策略性代理处获取诚实报告;尖锐阈值才是保持校准的设计。
原文摘要 · Abstract (English)
Eliciting truthful reports from autonomous agents is a core problem in scalable AI oversight: a principal scores the agent's report using a strictly proper scoring rule, but the agent also benefits from the report through a non-accuracy channel (approval for autonomous action, allocation share, downstream control). The same structure appears in classical mechanism-design settings such as marketplace operation. Our main result is an endogeneity: the principal's optimal oversight necessarily uses a non-affine approval function to screen types, yet any non-affine approval makes truthful reporting suboptimal under the combined objective whenever deviation is undetectable. The principal cannot avoid the perturbation that undermines calibration. This impossibility holds for all strictly proper scoring rules, with a closed-form perturbation formula. A constructive escape exists: a step-function approval threshold achieves first-best screening for every strictly proper scoring rule, because the agent's binary inflate-or-not choice creates a type-space threshold regardless of the generator's curvature. Under the Brier score specifically, the type-independent inflation cost yields a welfare equivalence between second-best and first-best; we prove this equivalence is unique to Brier (the welfare gap under smooth $C^1$ oversight is bounded below by $Ω(\text{Var}(1/G'') (γ/β)^2)$ for every non-Brier rule). Two instances develop the framework: AI agent oversight (the lead motivating setting) and marketplace operation (a parallel mechanism-design domain). The message for AI alignment is direct: smooth scoring-based oversight cannot elicit truthful reports from a strategic agent; sharp thresholds are the calibration-preserving design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。