让AI评价器的判断有据可查,避免错误推理误导决策。
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

- 构建八维反事实判断立方体,追踪评价变化的依据来源。
- 提出判断凭证概念,用最小替换集复现评价结果,解释判断转变。
- 发现高准确率模型在复杂推理中严重失真,需分开报告准确率与一致性。
评价系统常因错误推理得出正确标签,这对依赖其决策的智能体构成重大风险。现有评估仅验证最终标签正确性,忽视判断变化是否基于有效证据、一致规则或合理适用性。本文通过基础、规范和权威三类来源,形式化评价推理的责任机制,构建八维反事实判断立方体以刻画判断更新。定义判断凭证为能复现修正结论的最小来源替换集合。推导黑箱评价器的认证成本边界,并提出ReasonBench基准,包含19,520个案例与7,200个对照。在冻结评估中,Qwen3-1.7B达到98.41%凭证准确率,立方体预测得分为96.99%,较前者稳定低1.42分,经Qwen3-0.6B复现验证。标准高准确率掩盖严重鲁棒性缺陷:语义保持的来源置换使有效凭证恢复率降至54.8%与49.2%。单源训练模型保持93.75%判决准确率,但对多源更新仅恢复7.16%凭证。置换重训练虽提升一致性至96.6%,却加剧立方体预测缺陷。结构化反事实监督无法确保鲁棒推理。研究强调,意识原因的评估必须分离预测与认证,同时报告转化一致性与标准准确率,以实现可信评价审计。
原文摘要 · Abstract (English)
Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or supplying training feedback. Standard evaluation only verifies final label correctness, ignoring whether judgment changes stem from valid evidence, consistent rules, or proper rule applicability. We formalize evaluator reasoning accountability via three core sources: grounds, norms, and authority. Varying these sources yields an eight-cell counterfactual judgment cube to characterize judgment updates. We define judgment receipts as minimal source replacement sets that reproduce revised verdicts to explain judgment transitions. We derive certification cost bounds for black-box evaluators and present ReasonBench, a policy and logical reasoning benchmark with verifiable receipts covering 19,520 cases and 7,200 controls. In frozen evaluations, Qwen3-1.7B reaches 98.41% receipt accuracy, while cube prediction scores 96.99%, a consistent 1.42-point drop validated by Qwen3-0.6B replication. Strong standard accuracy masks severe robustness flaws. Meaning-preserving source permutations reduce valid receipt recovery to 54.8% and 49.2% for direct and cube prediction. Models trained on simple single-source changes retain 93.75% verdict accuracy but recover only 7.16% of receipts for complex multi-source updates. Permutation retraining boosts consistency to 96.6% yet worsens cube prediction deficits. Structured counterfactual supervision fails to guarantee robust reasoning. We show reason-aware evaluation must decouple prediction and certification, reporting transformation consistency alongside standard accuracy for trustworthy evaluator auditing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。