提出部署完备性评估,检验基准测试能否真实决定实际部署决策。
Deployment-complete benchmarking
- 定义基准测试的部署完备性:证据与部署行为是否一致
- 97.9% 的 Tox21 数据纤维存在混合决策,暴露部署信息缺失
- 建议报告证据、行动、模糊度和完成成本,而非仅输出分数
基准测试在部署、采购和科学筛选中日益重要,但分数仅反映响应结果,未必对应实际部署动作。本文提出部署完备性基准测试,判断基准证据是否足以决定部署行为。当每个证据纤维内部署动作恒定时,基准即为完备;混合纤维暴露部署信息缺失,完成曲线量化解决模糊所需证据量。在受控响应空间中,基准通道的符合覆盖率为94.98%,但在未测量部署通道中降至10.07%;而响应排名区间达到94.91%覆盖。即使基准误差为零,最大残差规模下也仅能确认45.4%的候选者。公开审计显示不完备现象普遍:Tox21有97.9%纤维为混合状态,主版Matbench与JARVIS审计中认证分数中位数为零。在保留回放中,'先认证再获取'策略将Tox21误判率从1.19%降至0.027%,JARVIS从20.3%降至0.128%,同时可改变模型选择并识别部署相关探针。部署可用的基准应报告证据、支持行动、模糊度及完成成本,而非仅提供分数。
原文摘要 · Abstract (English)
Benchmarks increasingly guide deployment, procurement and scientific screening, yet a score supports only the response it records, not necessarily the deployment action. We introduce deployment-complete benchmarking, which tests whether benchmark evidence determines a deployment action. A benchmark is complete for a claim exactly when the action is constant on each evidence fiber; mixed fibers expose missing deployment information, and completion curves quantify the evidence required to resolve ambiguity. In controlled response spaces, benchmark-channel conformal coverage of 94.98% transferred poorly to an unmeasured deployment channel (10.07%), whereas response-rank intervals achieved 94.91% coverage; even zero benchmark error certified only 45.4% of candidates at the largest residual size. Public audits revealed incompleteness, including 97.9% mixed Tox21 fibers and zero median certifiable fraction in main Matbench and JARVIS audits. In held-out replays, certify-then-acquire reduced false decisions from 1.19% to 0.027% in Tox21 and from 20.3% to 0.128% in JARVIS, while changing model choice and identifying deployment-relevant probes. Deployment-ready benchmarks should report evidence, supported actions, ambiguity and completion cost rather than scores alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。