arXiv:2509.00063physics.chem-phcs.AI2025-09被引 1

构建化学错误检测与修正基准,评估大模型在分子描述中的可靠性。

MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Revision

  • 分四步评估:识别、定位、解释、修正分子描述中的化学错误。
  • 包含1193个细粒度标注实例,每条含错误类型、位置、解释和修正方案。
  • 适合关注化学大模型可信性与推理能力的研究者使用。

大语言模型在分子科学领域展现出巨大潜力,但常产生化学不准确的描述,且难以识别或解释潜在错误,引发对其在科研应用中稳健性和可靠性的担忧。为支持对大模型化学推理能力的更严格评估,本文提出MolErr2Fix基准,专门评测模型在分子描述中的错误检测与修正能力。不同于聚焦分子到文本生成或性质预测的现有基准,MolErr2Fix强调细粒度化学理解,要求模型识别、定位、解释并修正分子描述中的结构或语义错误。该基准包含1193个精细标注的错误实例,每个实例均具备四元标注:(错误类型、错误位置、解释理由、修正结果)。这些任务模拟了真实化学交流中所需的推理与验证过程。对当前主流大模型的评估揭示显著性能差距,凸显提升化学推理能力的迫切需求。MolErr2Fix旨在为相关能力提供专注评测工具,并推动更可靠、化学感知更强的语言模型发展。所有标注数据及配套评估API将公开发布,以促进后续研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown growing potential in molecular sciences, but they often produce chemically inaccurate descriptions and struggle to recognize or justify potential errors. This raises important concerns about their robustness and reliability in scientific applications. To support more rigorous evaluation of LLMs in chemical reasoning, we present the MolErr2Fix benchmark, designed to assess LLMs on error detection and correction in molecular descriptions. Unlike existing benchmarks focused on molecule-to-text generation or property prediction, MolErr2Fix emphasizes fine-grained chemical understanding. It tasks LLMs with identifying, localizing, explaining, and revising potential structural and semantic errors in molecular descriptions. Specifically, MolErr2Fix consists of 1,193 fine-grained annotated error instances. Each instance contains quadruple annotations, i.e,. (error type, span location, the explanation, and the correction). These tasks are intended to reflect the types of reasoning and verification required in real-world chemical communication. Evaluations of current state-of-the-art LLMs reveal notable performance gaps, underscoring the need for more robust chemical reasoning capabilities. MolErr2Fix provides a focused benchmark for evaluating such capabilities and aims to support progress toward more reliable and chemically informed language models. All annotations and an accompanying evaluation API will be publicly released to facilitate future research.

化学大模型错误修正可信性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。