让大模型化学推理过程可验证,发现答案对但逻辑错的隐藏问题
From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

- 用规则引擎自动检查化学推理每一步是否合法
- 5620个样本覆盖4类任务,区分答案对与逻辑对
- 适合评估化学AI助手真实可靠性,不依赖人工打分
大语言模型在化学领域日益作为助手使用,但现有评测仅关注最终答案,掩盖了关键缺陷:模型可能输出正确分子或产物,但其推理过程违背化学逻辑。现有过程级评估成本高、不一致且易幻觉。我们提出 ChemCoTBench-V2,一个基于规则可验证的诊断基准,支持低成本、可审计的结构化推理评估。涵盖分子理解、编辑、优化和反应预测,共5,620个样本,18项任务。模型需按专家设计模板暴露关键中间步骤,这些步骤通过确定性化学规则及参考轨迹(闭合答案)进行验证,而非依赖另一个LLM判别;开放任务则用可验证状态约束评估。基准报告三项信号:最终答案正确性、模板遵循度、以及专家精炼中间承诺的步骤级验证正确率。前沿模型实验显示,最终答案正确与结构化推理一致性间存在持续差距:模型常格式合规但化学步骤失败,或答案正确但支撑推理薄弱。该基准支持细粒度对比,并定位首次违反验证的步骤。
原文摘要 · Abstract (English)
Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing process-level evaluators are hard to scale because LLM judges and human step-level process annotation are costly, inconsistent, and vulnerable to hallucination. We introduce ChemCoTBench-V2, a rule-verifiable diagnostic benchmark for low-cost, auditable evaluation of structured, verifier-addressable chemical reasoning traces. It spans molecular understanding, molecule editing, molecular optimization, and reaction prediction, with 5,620 evaluation samples across 18 reporting tasks. Models must expose key intermediate steps in expert-designed templates, and those steps are checked with deterministic chemistry rules and, for closed-answer tasks, reference traces rather than another LLM judge. Open-ended molecular optimization is evaluated with oracle-verifiable state constraints rather than strict trace matching. The benchmark reports three separate signals: final-answer correctness, template adherence, and step-wise verifier correctness over expert-refined intermediate commitments. Experiments on frontier models reveal a persistent gap between final-answer success and structured-reasoning-state consistency: models often follow the requested format while failing chemical-step checks, or answer correctly with weak supporting reasoning. ChemCoTBench-V2 enables fine-grained model comparison and identifies the concrete step at which the trace first violates the verifier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。