arXiv:2508.01670cs.AIphysics.chem-ph2025-08被引 7

测试大模型在化学计算中的推理能力,发现越复杂任务表现越差。

QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry

  • 构建7个化学领域350道计算题的评测基准
  • 模型在难题上准确率显著下降,显示计算短板
  • 适合关注科学推理与LLM可靠性的研究者

定量化学是现代化学研究的核心,但大语言模型(LLMs)在执行严谨、逐步计算方面的能力仍缺乏深入探索。为此,我们提出QCBench,一个面向定量化学的基准测试,涵盖7个化学子领域(分析化学、生物/有机化学、普通化学、无机化学、物理化学、高分子化学和量子化学)共350道计算题。为系统评估大模型的数学推理能力,题目分为简单、中等、困难三类。每道题基于真实化学场景设计,防止投机取巧,要求明确数值推理。QCBench可实现对计算弱点的细粒度诊断,揭示模型在不同难度下的局限性,并为未来改进(如领域自适应微调或多模态融合)提供基础。对24个LLM的评估显示,任务复杂度越高,性能越下降,凸显当前语言流畅性与科学计算准确性之间的差距。QCBench代码已开源:https://github.com/jiaqingxie/QCBench。

原文摘要 · Abstract (English)

Quantitative chemistry is central to modern chemical research, yet the ability of large language models (LLMs) to perform its rigorous, step-by-step calculations remains underexplored. To fill this blank, we propose QCBench, a Quantitative Chemistry oriented benchmark comprising 350 computational chemistry problems across 7 chemistry subfields, which contains analytical chemistry, bio/organic chemistry, general chemistry, inorganic chemistry, physical chemistry, polymer chemistry and quantum chemistry. To systematically evaluate the mathematical reasoning abilities of large language models (LLMs), they are categorized into three tiers: easy, medium, and difficult. Each problem, rooted in realistic chemical scenarios, is structured to prevent heuristic shortcuts and demand explicit numerical reasoning. QCBench enables fine-grained diagnosis of computational weaknesses, reveals model-specific limitations across difficulty levels, and lays the groundwork for future improvements such as domain-adaptive fine-tuning or multi-modal integration. Evaluations on 24 LLMs demonstrate a consistent performance degradation with increasing task complexity, highlighting the current gap between language fluency and scientific computation accuracy. Code for QCBench is available at https://github.com/jiaqingxie/QCBench.

化学推理大模型评测量化计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。