测试大模型用科学工具做化学模拟的可靠性,发现工具有用但也会出错。
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

- 构建200道基于真实地质化学场景的多选题,检验模型调用模拟器的能力。
- 使用工具后整体准确率提升,但部分原本正确的题反而答错,暴露隐藏缺陷。
- 输出查看方式影响效果:强模型用目录式接口更省资源,中等模型则易出错。
大型语言模型代理正越来越多地连接到科学软件,但尚不清楚何时工具接入能提升科学计算的可靠性,而非仅仅增加复杂性。我们提出PHREEQC-MCQ-200,一个用于评估工具增强型代理在确定性水相地球化学模拟任务中表现的基准。该基准包含200道多选题,源自21个经过验证的PHREEQC场景,要求代理构建模拟输入、执行PHREEQC、检查结构化输出,并提交最终答案。在多个前沿与中等水平模型家族中,使用模拟器显著提升了总体准确率,证实了在许多科学计算任务中,基于执行的推理是必要的。然而,提升并非单调:工具增强代理也出现了原本正确的问题反而答错的情况,揭示了平均准确率所掩盖的退化现象。我们进一步表明,输出访问协议至关重要。表格目录界面可在保持或提升强模型准确率的同时降低令牌开销,但对于无法可靠导航结构化输出的中等模型则会降低性能。因此,PHREEQC-MCQ-200将科学工具使用视为端到端诊断问题,而非简单的工具调用能力。我们认为,科学代理评估应不仅报告准确率,还应包括项目级保留率、输出访问敏感性、轨迹失败情况及计算链断裂位置。
原文摘要 · Abstract (English)
Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a benchmark for evaluating tool-augmented agents on deterministic aqueous-geochemistry simulations. The benchmark contains 200 multiple-choice questions derived from 21 validated PHREEQC scenarios, requiring agents to construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to final answers. Across multiple frontier and mid-tier model families, simulator access substantially improves aggregate accuracy, confirming that grounded execution is necessary for many scientific-computation tasks. However, the gains are not monotonic: tool-augmented agents also lose items they answered correctly without tools, revealing regressions that average accuracy alone hides. We further show that output-access protocol matters. A table-of-contents interface can reduce token cost while preserving or improving accuracy for stronger models, but it degrades performance for mid-tier models that cannot reliably navigate structured simulator outputs. PHREEQC-MCQ-200 therefore frames scientific tool use as an end-to-end diagnostic problem rather than a simple tool-calling capability. We argue that evaluations of scientific agents should report not only accuracy, but also item-level retention, output-access sensitivity, trajectory failures, and where the computation chain breaks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。