arXiv:2608.21504cs.LG2026-08

ChemDIRT评测化学大模型在多样指令与分子表示下的推理鲁棒性。

ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation

论文配图:ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation
图 1 · 摘自论文原文
  • 设计多维度评测框架,测试指令与分子表示变化下的表现
  • 发现模型对提示词和表示方式高度敏感,性能波动大
  • 适合评估化学大模型真实推理能力,尤其关注鲁棒性

大型语言模型(LLMs)在化学等科学领域的应用日益广泛,但现有化学评测基准通常仅涵盖有限任务,忽视问题表述和分子表示变化带来的影响,导致性能评估过于乐观。为此,我们提出ChemDIRT(多样化指令、表示与任务基准),一个全面评估化学推理鲁棒性的框架。该框架系统考察模型在指令和分子表示变化下的表现,覆盖八类化学任务。通过在受控扰动下评估准确率与一致性,ChemDIRT比传统单格式基准更可靠。我们对多种开源与闭源模型进行评测,发现显著的提示敏感性、表示依赖性及任务类别间性能不均。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry. However, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to variations in problem formulation and chemical representation. As a result, reported performance may overestimate a model's true ability to reason consistently across realistic settings. To address this challenge, we introduce ChemDIRT (Diversified Instruction, Representation, and Task Benchmark), a comprehensive evaluation framework designed to assess the robustness of chemical reasoning in LLMs. ChemDIRT systematically measures model performance across variations in instructions and molecular representations while spanning eight categories of chemistry tasks. By evaluating both accuracy and consistency under these controlled perturbations, ChemDIRT provides a more reliable assessment of model reasoning capabilities than conventional single-format benchmarks. We benchmark a diverse set of open- and closed-source LLMs, revealing substantial prompt sensitivity, representation dependence, and uneven performance across task families.

化学大模型评测基准鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。