arXiv:2602.03108cs.CL2026-02中稿 · Artificial Intelli…被引 1

ChemPro为大模型设计了渐进式化学评测,测试从基础到高阶的科学推理能力。

ChemPro: A Progressive Chemistry Benchmark for Large Language Models

  • 按难度分4个阶段,覆盖有机、无机、物化等多领域化学题型。
  • 45+模型测试显示:基础题表现尚可,复杂题准确率显著下降。
  • 适合评估大模型科学理解力,推动更鲁棒的推理方法研究。

我们提出ChemPro,一个包含4100个自然语言问答对的渐进式化学基准,涵盖4个连贯难度等级,用于评估大语言模型(LLMs)在广泛化学主题中的表现。题目包括选择题和数值题,覆盖精细信息回忆、长程推理、多概念综合、需细致表达的问题求解及基础题,比例均衡,全面覆盖生物化学、无机化学、有机化学和物理化学。ChemPro设计类比学生从中学到高中化学的学习评估过程,难度逐步提升,严格检验模型从解决基础问题到应对复杂挑战的能力。我们评估了45+7个先进大模型,涵盖开源与专有版本,分析发现:尽管模型在基础题上表现良好,但在不同复杂度类型下准确率持续下降。结果揭示了大模型在通用科学推理与理解方面的关键局限,指出尚未充分研究的难度维度,强调需要更稳健的方法来提升模型能力。

原文摘要 · Abstract (English)

We introduce ChemPro, a progressive benchmark with 4100 natural language question-answer pairs in Chemistry, across 4 coherent sections of difficulty designed to assess the proficiency of Large Language Models (LLMs) in a broad spectrum of general chemistry topics. We include Multiple Choice Questions and Numerical Questions spread across fine-grained information recall, long-horizon reasoning, multi-concept questions, problem-solving with nuanced articulation, and straightforward questions in a balanced ratio, effectively covering Bio-Chemistry, Inorganic-Chemistry, Organic-Chemistry and Physical-Chemistry. ChemPro is carefully designed analogous to a student's academic evaluation for basic to high-school chemistry. A gradual increase in the question difficulty rigorously tests the ability of LLMs to progress from solving basic problems to solving more sophisticated challenges. We evaluate 45+7 state-of-the-art LLMs, spanning both open-source and proprietary variants, and our analysis reveals that while LLMs perform well on basic chemistry questions, their accuracy declines with different types and levels of complexity. These findings highlight the critical limitations of LLMs in general scientific reasoning and understanding and point towards understudied dimensions of difficulty, emphasizing the need for more robust methodologies to improve LLMs.

化学大模型评测基准科学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。