SUPERChem为化学大模型提供多模态推理评测基准,挑战专家级化学思维能力。
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
- 构建500道专家设计的多步化学推理题,支持图文双格式。
- 模型最高准确率38.5%,远低于人类40.3%的基准表现。
- 引入推理路径一致性评分,可区分真实推理与表面匹配。
当前评估大语言模型(LLM)化学推理能力的基准存在任务过于简单、缺乏过程评估和与专家技能脱节等问题。为此,我们推出SUPERChem,一个包含500道专家精心设计的高难度化学推理题的基准,覆盖多个子领域,并提供多模态与纯文本两种形式。通过原创内容和迭代式审核流程,剔除错误题目并减少数据污染。每道题均配有专家撰写的解题路径,支持推理路径保真度(RPF)评分,可超越最终答案准确率,评估推理质量。在与人类基准(40.3%准确率)对比中,表现最好的模型GPT-5(High)仅达38.5%,其次为Gemini 2.5 Pro(37.9%)和DeepSeek-V3.1-Think(37.3%)。SUPERChem能激发多步、多模态推理,揭示视觉信息对模型的影响差异,并区分高质量推理与启发式猜测。该基准提供挑战性测试环境与可靠评估框架,助力大模型向专家级化学智能演进。数据集已公开于https://huggingface.co/datasets/ZehuaZhao/SUPERChem。
原文摘要 · Abstract (English)
Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and misalignment with expert-level chemistry skills. To address these issues, we introduce SUPERChem, a benchmark of 500 expert-curated reasoning-intensive chemistry problems, covering diverse subfields and provided in both multimodal and text-only formats. Original content and an iterative curation pipeline eliminate flawed items and mitigate data contamination. Each problem is paired with an expert-authored solution path, enabling Reasoning Path Fidelity (RPF) scoring to evaluate reasoning quality beyond final-answer accuracy. Evaluations against a human baseline of 40.3% accuracy show that even the best-performing model, GPT-5 (High), reaches only 38.5%, followed closely by Gemini 2.5 Pro (37.9%) and DeepSeek-V3.1-Think (37.3%). SUPERChem elicits multi-step, multimodal reasoning, reveals model-dependent effects of visual information, and distinguishes high-fidelity reasoners from heuristic ones. By providing a challenging benchmark and a reliable evaluation framework, SUPERChem aims to facilitate the advancement of LLMs toward expert-level chemical intelligence. The dataset of the benchmark is available at https://huggingface.co/datasets/ZehuaZhao/SUPERChem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。