arXiv:2605.25997cs.LGstat.ML2026-05被引 1

提出部署完备性评估,检验基准测试能否真实决定实际部署决策。

Deployment-complete benchmarking

  • 定义基准测试的部署完备性:证据与部署行为是否一致
  • 97.9% 的 Tox21 数据纤维存在混合决策,暴露部署信息缺失
  • 建议报告证据、行动、模糊度和完成成本,而非仅输出分数

基准测试在部署、采购和科学筛选中日益重要,但分数仅反映响应结果,未必对应实际部署动作。本文提出部署完备性基准测试,判断基准证据是否足以决定部署行为。当每个证据纤维内部署动作恒定时,基准即为完备;混合纤维暴露部署信息缺失,完成曲线量化解决模糊所需证据量。在受控响应空间中,基准通道的符合覆盖率为94.98%,但在未测量部署通道中降至10.07%;而响应排名区间达到94.91%覆盖。即使基准误差为零,最大残差规模下也仅能确认45.4%的候选者。公开审计显示不完备现象普遍:Tox21有97.9%纤维为混合状态,主版Matbench与JARVIS审计中认证分数中位数为零。在保留回放中,'先认证再获取'策略将Tox21误判率从1.19%降至0.027%,JARVIS从20.3%降至0.128%,同时可改变模型选择并识别部署相关探针。部署可用的基准应报告证据、支持行动、模糊度及完成成本,而非仅提供分数。

原文摘要 · Abstract (English)

Benchmarks increasingly guide deployment, procurement and scientific screening, yet a score supports only the response it records, not necessarily the deployment action. We introduce deployment-complete benchmarking, which tests whether benchmark evidence determines a deployment action. A benchmark is complete for a claim exactly when the action is constant on each evidence fiber; mixed fibers expose missing deployment information, and completion curves quantify the evidence required to resolve ambiguity. In controlled response spaces, benchmark-channel conformal coverage of 94.98% transferred poorly to an unmeasured deployment channel (10.07%), whereas response-rank intervals achieved 94.91% coverage; even zero benchmark error certified only 45.4% of candidates at the largest residual size. Public audits revealed incompleteness, including 97.9% mixed Tox21 fibers and zero median certifiable fraction in main Matbench and JARVIS audits. In held-out replays, certify-then-acquire reduced false decisions from 1.19% to 0.027% in Tox21 and from 20.3% to 0.128% in JARVIS, while changing model choice and identifying deployment-relevant probes. Deployment-ready benchmarks should report evidence, supported actions, ambiguity and completion cost rather than scores alone.

基准测试部署验证机器学习可信性模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。