arXiv:2603.25780cs.SEcs.LG2026-03被引 1

用智能裁判自动验证科学模拟代码,误判率从42%降到1.5%

A Judge Agent Closes the Reliability Gap in AI-Generated Scientific Simulation

  • 设计裁判代理自动化验证数学正确性
  • 在134个跨领域测试中失败率降至1.5%
  • 适合需高可靠性的科研与临床模拟场景

大语言模型可生成科学模拟代码,但在非教材问题上常无声失败。本文提出裁判代理(Judge Agent),可全自动执行经典数学验证——适定性、收敛性与误差认证,在134个跨越12个科学领域的测试案例中,将无声失败率从42%降低至1.5%。前瞻性基准测试中,72个由12位独立科学家盲测的任务,使用自动误差界时成功率达89%(95%置信区间:[80%, 95%]),无裁判时仅为53%。在临床CT(唯一实证实验,n=200)中,该流程达到专家水平的99%。剩余1.5%失败集中在分岔点,此处可证性失效。论文定义了可模拟类S,并提出spec.md结构化规范格式,使科学计算问题机器可读且求解器无关。代码、数据及全部72个基准任务均已公开归档。

原文摘要 · Abstract (English)

Large language models can generate scientific simulation code, but the generated code silently fails on most non-textbook problems. We show that classical mathematical validation -- well-posedness, convergence, and error certification -- can be fully automated by a Judge Agent, reducing the silent-failure rate from 42% to 1.5% across 134 test cases spanning 12 scientific domains. The headline result comes from a prospective benchmark: 72 blinded tasks submitted by 12 independent scientists yield an 89% success rate (95% CI: [80%, 95%]) with automated error bounds, versus 53% without the Judge. On clinical CT (the only powered experiment, n = 200), the pipeline reaches 99% of expert quality. The residual 1.5% concentrates at bifurcation points where certifiability breaks down. We formalize this boundary through the simulability class S and introduce spec.md, a structured specification format that makes any scientific computation problem machine-readable and solver-independent. Code, data, and all 72 benchmark tasks are publicly archived.

AI验证科学模拟可靠性裁判代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。