测试大模型在化学反应成本估算中的实际能力,发现工具调用仍不足以准确完成任务。
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning

- 构建化学采购成本基准,需精准识别化学品、查价、选包装并计算总价
- 顶尖模型在干净输入下仅达50.6%准确率,噪声环境下性能大幅下降
- 适合关注科学智能体评估与化学领域应用的科研人员参考
大型语言模型作为工具调用智能体的能力日益提升,但其在科学任务中的严谨评估仍显不足。本文聚焦化学合成中的采购成本估算,该任务要求智能体准确识别化学品、检索供应商报价、选择可购包装、统一数量单位并计算总成本。我们提出ChemCost基准,包含1,427个可评估反应,基于2,261种化学品和230,775条供应商报价的静态快照,支持对解析、检索、采购及算术等环节的逐阶段诊断。为测试鲁棒性,还构建了引入化学别名扰动、数量表达错误、字段缺失和格式混乱的噪声版本。实验表明,工具访问虽必要但不充分,最强模型在干净输入下仅50.6%准确率且相对误差在25%以内,噪声下性能显著下降。阶段分析揭示失败源于脆弱的解析、无效的信息整合、无效包装选择及非收敛的工具使用。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become increasingly capable as tool-using agents, with benchmarks spanning diverse general agentic tasks. Yet rigorous evaluation of scientific tool use remains limited. In chemistry, recent agents can plan syntheses and invoke domain-specific tools, but evaluations often rely on curated demonstrations, expert assessment, or LLM-as-judge scoring rather than exact, judge-free ground truth. We address this gap with chemical procurement cost estimation, a practical task in which an agent must ground chemical identities, retrieve supplier quotes, select valid purchasable packs, normalize quantities, and compute cost from a reaction description. We introduce ChemCost, a benchmark of 1,427 evaluable reactions grounded to a frozen pricing snapshot covering 2,261 chemicals and 230,775 supplier quotes, supporting scalar scoring and stage-level diagnosis of grounding, retrieval, procurement, and arithmetic failures. To evaluate robustness, we further construct controlled noise-injected views that perturb chemical aliases, quantity expressions, missing fields, and input formatting. Experiments with frontier, open-weight, and chemistry-specialized LLM agents show that tool access is necessary but insufficient for solving the task. The strongest agents reach only 50.6% accuracy within 25% relative error on clean inputs and degrade substantially with realistic noise. Stage-level analysis further shows that failures arise from brittle parsing, ineffective evidence integration, invalid pack selection, and non-convergent tool use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。