用可执行评估暴露大模型在环境科学计算中的隐藏错误
Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science

- 构建可运行的计算过程评测框架,让每一步推理可见
- 多选题使准确率虚高至少12个百分点,真实能力被夸大
- 顶尖模型遇新条件仍不会灵活调整方法,需专家把关
大型语言模型在环境科学定量任务中日益普及,但现有评估仅关注最终答案,忽略计算过程。本文提出AtmosCoder-Bench,一个基于可执行验证的基准测试,通过可迁移的半自动化流程构建(436个问题,3,910种变体,7,029个评分量),确保每个问题无歧义且人类可解,答案唯一可验证。实验发现:(i) 多选题形式使测得准确率至少虚高12个百分点;(ii) 许多失败并非因知识缺失,而是模型在多步计算中无法一致应用已知公式与约束;(iii) 即便前沿模型,在任务特定条件使常用方法失效时,仍常退回到经典解法模式,而非适配实际物理机制,凸显专家监督的必要性。
原文摘要 · Abstract (English)
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an execution-grounded benchmark that makes the calculation process visible. Built through a transferable semi-automated pipeline (436 problems, 3,910 variants, 7,029 graded quantities), every problem is validated to be unambiguous and human-solvable, with uniquely verifiable answers. We find that (i) multiple-choice formats inflate measured accuracy by at least 12 percentage points; (ii) many failures arise not from missing knowledge but from models failing to apply known formulas and constraints consistently throughout multi-step computation; and (iii) even frontier models remain weak when task-specific conditions invalidate familiar methods, often reverting to canonical solution patterns rather than adapting methods to the relevant physical regime, leaving expert oversight essential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。