arXiv:2608.01378cs.LG2026-08

模型能否替代实验?关键看是否通过针对性审计,否则再准也会选错。

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

论文配图:When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design
图 1 · 摘自论文原文
  • 用审计代替盲目信任,确保模型推荐结果真实可靠
  • 筛选候选会放大预测误差,形成可量化的‘选择税’
  • 仅凭准确率不可信,需通过审计降低评估成本25倍

化学、材料科学和机器学习中的设计流程常受限于昂贵的评估:实验、第一性原理模拟或完整训练。越来越多地使用机器学习代理模型来预测这些结果,并用于候选推荐、评分,甚至将自身预测当作测量值反馈回搜索过程。通过对三个全面验证的任务进行数学分析,我们确定了该做法的安全边界、安全证书的成本,以及替代是否真正划算。预测精度无法支撑信任:接近完美的R²仍可能导致最差选择;筛选N个候选会带来可量化的‘选择税’,其上下界均能计算。安全的关键在于架构规则——模型可自由提议和训练,但所有认证结论必须基于真实评估。这一规则在无任何假设下充分且必要,因若允许预测以测量身份进入认证,将引发确定性的自我确认失效。我们推导出模型可充当预言机的最小条件(排名保真,非精度),证明信任必须通过选择感知的审计来购买,且审计在查询复杂度上最优。在6种任务条件下432次代理拟合中,审计统计量与实际搜索表现的斯皮尔曼相关系数达0.80–0.99,而R²与实际遗憾的相关性低至0.33;经审计筛选后,认证预言机成本降低25倍。

原文摘要 · Abstract (English)

Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run. Machine-learning surrogates that predict these outcomes are increasingly used not only to propose candidates but to grade them, and even to feed their own predictions back into the search as though they were measurements. Through mathematical analysis validated on three exhaustively ground-truthed design tasks, we establish when this practice is safe, what any certificate of safety must cost, and when the substitution provably pays. Predictive accuracy cannot anchor trust: near-perfect R^2 is compatible with worst-possible selections, and screening N candidates inflates the over-prediction at the selected candidate by a quantifiable "selection tax" with matching upper and lower bounds. Safety follows instead from an architectural rule - predictions may propose and train without restriction, but every certified conclusion must rest on true evaluations - which is sufficient with no assumptions on the surrogate, and necessary, since admitting predictions into certification with the standing of measurements opens a deterministic self-confirmation failure mode. We derive the minimal criterion under which a model may act as an oracle (rank preservation, not accuracy), show that trust must be purchased through selection-aware audits that are optimal in query complexity, and prove a dichotomy fixing when audited surrogates cut certified evaluation cost. Across 432 surrogate fits over six task-regime conditions, the audit statistic tracks deployed search performance at Spearman rank correlation 0.80-0.99, while the rank correlation of R^2 with deployed regret falls as low as 0.33; audited screening reduces certified oracle cost by a measured factor of 25.

模型可信度代理模型审计机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。