提出新评估方法,让AI内容审核不再被历史标签误导。
Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI

- 用规则可推导性替代人工标签一致性评估
- 发现传统方法误判率高达79.8%,实际是规则模糊所致
- 适合需要高合规性的AI治理系统开发者
内容审核系统通常依赖与人工标注的一致性进行评估。但在规则约束环境中,这一假设失效:多个决策可能均符合政策逻辑,而一致性的度量会惩罚合法判断,将规则模糊误判为错误——我们称之为‘协议陷阱’。本文将评估定义为基于政策的正确性,引入可辩护性指数(DI)和模糊性指数(AI)。为避免额外审计开销,提出基于审计模型词元对数概率的随机可辩护信号(PDS),将大模型推理轨迹作为治理信号,而非分类输出:审计模型不直接决定是否违规,而是验证某决策是否可从规则层级中逻辑推导。在超过19.3万条Reddit社区审核决策上验证,基于一致性与基于政策的评估存在33-46.6个百分点的差距,79.8%-80.6%的模型假阴性实为政策合理决策。进一步发现,规则细化可使模糊性指数下降10.8个百分点,而可辩护性稳定。重复采样分析表明PDS方差主要源于治理模糊性而非解码噪声。基于这些信号构建的治理门控系统实现78.6%自动化覆盖率,风险降低64.9%。结果表明,规则约束环境下的评估应从匹配历史标签转向基于规则的推理有效性。
原文摘要 · Abstract (English)
Content moderation systems are typically evaluated by measuring agreement with human labels. In rule-governed environments this assumption fails: multiple decisions may be logically consistent with the governing policy, and agreement metrics penalize valid decisions while mischaracterizing ambiguity as error -- a failure mode we term the Agreement Trap. We formalize evaluation as policy-grounded correctness and introduce the Defensibility Index (DI) and Ambiguity Index (AI). To estimate reasoning stability without additional audit passes, we introduce the Probabilistic Defensibility Signal (PDS), derived from audit-model token logprobs. We harness LLM reasoning traces as a governance signal rather than a classification output by deploying the audit model not to decide whether content violates policy, but to verify whether a proposed decision is logically derivable from the governing rule hierarchy. We validate the framework on 193,000+ Reddit moderation decisions across multiple communities and evaluation cohorts, finding a 33-46.6 percentage-point gap between agreement-based and policy-grounded metrics, with 79.8-80.6% of the model's false negatives corresponding to policy-grounded decisions rather than true errors. We further show that measured ambiguity is driven by rule specificity: auditing 37,286 identical decisions under three tiers of the same community rules reduces AI by 10.8 pp while DI remains stable. Repeated-sampling analysis attributes PDS variance primarily to governance ambiguity rather than decoding noise. A Governance Gate built on these signals achieves 78.6% automation coverage with 64.9% risk reduction. Together, these results show that evaluation in rule-governed environments should shift from agreement with historical labels to reasoning-grounded validity under explicit rules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。