arXiv:2608.19269cs.SEcs.AI2026-08
冻结评估数据集,验证历史结论可复现性,揭示评估结果的可信边界。
What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals
- 冻结大规模评估集,尝试复现历史结论
- 多数结论因证据未绑定而无法复现
- 可复现结论在不同精度下保持稳定
评估基准在不明确其结果所许可范围的情况下运行。我们冻结了一个大型评估集合,并尝试复现其历史主张。大多数评估单元因复现所需证据未被绑定而无法继续。当复现可行时,不同主张在不同分辨率下保持稳定。本工作将这一原本隐含的推理步骤显式化并实现可执行。
原文摘要 · Abstract (English)
Benchmarks can run without determining what their results license. We freeze a large evaluation collection and attempt to replay its historical claims. Most units stop because the evidence required for replay is not bound. Where replay is possible, different claims remain stable at different resolutions. We make this otherwise implicit inference step explicit and executable.
评估基准可复现性证据绑定
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。