评测大模型在化学文献中的反应理解能力,发现现有模型仍难精准识别反应逻辑。
RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature
- 构建双任务基准:单图问答与全文问答,覆盖真实文献中的图文信息。
- 模型在全文问答中最高仅达49%准确率,显式文本提取优于深层推理。
- 强调需专用视觉编码器和更强推理机制,适合研发智能化学AI的研究者。
将多模态大语言模型(MLLMs)融入化学领域有望推动科学发现,但其对真实文献中密集、图形化反应描述的理解能力尚未充分探索。本文提出RxnBench,一个多层次基准,用于严格评估MLLMs在化学反应理解方面的表现。该基准包含两项任务:单图问答(SF-QA),基于305个精选反应图谱生成1,525个问题,测试细粒度视觉感知与机理推理;全文档问答(FD-QA),要求模型从108篇论文中整合文本、图示和表格信息,实现跨模态理解。对MLLMs的评估显示显著能力差距:尽管模型在提取显式文本方面表现良好,但在深层化学逻辑与精确结构识别上表现不佳。值得注意的是,具有推理时推理能力的模型显著优于标准架构,但所有模型在FD-QA上的准确率均未超过50%。这些结果凸显了开发领域专用视觉编码器和更强推理引擎的迫切需求,以推进自主化学人工智能的发展。
原文摘要 · Abstract (English)
The integration of Multimodal Large Language Models (MLLMs) into chemistry promises to revolutionize scientific discovery, yet their ability to comprehend the dense, graphical language of reactions within authentic literature remains underexplored. Here, we introduce RxnBench, a multi-tiered benchmark designed to rigorously evaluate MLLMs on chemical reaction understanding from scientific PDFs. RxnBench comprises two tasks: Single-Figure QA (SF-QA), which tests fine-grained visual perception and mechanistic reasoning using 1,525 questions derived from 305 curated reaction schemes, and Full-Document QA (FD-QA), which challenges models to synthesize information from 108 articles, requiring cross-modal integration of text, schemes, and tables. Our evaluation of MLLMs reveals a critical capability gap: while models excel at extracting explicit text, they struggle with deep chemical logic and precise structural recognition. Notably, models with inference-time reasoning significantly outperform standard architectures, yet none achieve 50\% accuracy on FD-QA. These findings underscore the urgent need for domain-specific visual encoders and stronger reasoning engines to advance autonomous AI chemists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。