arXiv:2609.01992cs.AIcs.CR2026-09

提出可验证证据充分性与覆盖度的ClaimReceipt系统,确保智能体评估结果可信。

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

  • 为每个主张生成绑定实验清单的可验证收据,支持精确审计
  • 在1392条记录中复现所有人工标注结果,错误率0且无误报
  • 适合需要高可信度验证的AI实验评估场景

智能体评估面临两个证据问题:报告主张能否从保留证据中复现(充分性),以及保留记录是否覆盖全部实验(覆盖度)。通用日志和哈希链无法可靠回答。我们提出ClaimReceipt,一种基于主张的收据规范与选择性验证机制,将类型化交易证据绑定至签名实验清单,并对每项主张返回通过、无效或不确定结果。规范在实现前冻结(SHA-256 18d109...b81)。在1,392条历史买卖记录上,CR-2验证器准确重现全部5个手动标注的审计结论,完全重放600条确定性及792条后生成记录,测试中13个声明字段组均无冗余,11个语义故障中11个返回预期结果,0个误报。随后进行独立前瞻性实验(CR-3):30项任务提前承诺,终端收据签名并链式连接,私有证据加密供审计。完整证据获得覆盖与会计通过;仅缺一个终端收据时返回不确定覆盖,而所有私有开启被隐藏则保持覆盖与协议验证通过,但经济主张无法判断,精准匹配预注册预测。收据注入仅增加0.021%模型推理时间,每笔交易增加9.9KB。规范可读性探测表明,我们冻结的规范对独立读者仍不够清晰。因此,主张验证需同时具备足够证据与承诺的实验宇宙以暴露遗漏。

原文摘要 · Abstract (English)

Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked transcripts answer neither reliably. We introduce ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim. We freeze the specification before implementation (SHA-256 18d109...b81). On 1,392 historical buyer--seller records, a CR-2 verifier reproduces all five manually labeled audit verdicts, exactly replays 600 deterministic and 792 post-generation records, makes every one of 13 declared field groups non-redundant under tested ablations, and returns the expected result on 11/11 semantic faults with 0/8 false positives. We then run a separate prospective CR-3 epoch: 30 assignments are committed before inference, terminal receipts are signed and chained, and private evidence is encrypted for an auditor. Complete evidence yields coverage and accounting PASS; withholding one terminal receipt returns INCONCLUSIVE_COVERAGE, while withholding all private openings preserves coverage and protocol verification but makes economic claims inconclusive, exactly matching a preregistered prediction. Receipt instrumentation adds 0.021% of model-inference time and 9.9 KB per transaction. A specification-legibility probe indicates that our own frozen specification is not yet unambiguous to an independent reader. Claim verification therefore requires both claim-sufficient evidence and a committed universe against which omissions become visible.

可信评估证据验证智能体审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。