构建化学表格多模态评测基准,揭示模型在科学理解上的短板
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
- 基于真实文献构建大规模化学表格数据集,含专家标注的布局与语义
- 主流模型在表格结构识别上表现尚可,但对分子结构等符号理解仍严重不足
- 适合研究科学智能、多模态推理及领域专用模型的开发者使用
随着多模态大模型在科学智能中的广泛应用,亟需更具挑战性的评估基准来检验其对复杂科学数据的理解能力。科学表格作为知识表征的核心载体,融合了文本、符号与图形,构成典型的多模态推理场景。然而现有基准多聚焦通用领域,难以反映科学研究中特有的结构复杂性与领域语义。化学表格尤为典型:其将反应物、条件、产率等结构化变量与分子结构、化学式等视觉符号交织在一起,对跨模态对齐与语义解析提出严峻挑战。为此,我们提出了ChemTable——一个源自真实文献的大规模化学表格基准,包含专家标注的单元格布局、逻辑结构及领域特定标签。该基准支持两项核心任务:(1) 表格识别(结构与内容提取);(2) 表格理解(描述性与推理型问答)。在ChemTable上的评估显示,尽管主流多模态模型在布局解析方面表现合理,但在处理分子结构与符号规范等关键元素时仍存在显著局限。闭源模型整体领先,但仍远未达到人类水平。本工作为科学多模态理解提供了真实测试平台,揭示了当前领域专用推理的瓶颈,推动科学智能系统的进步。
原文摘要 · Abstract (English)
With the widespread application of multimodal large language models in scientific intelligence, there is an urgent need for more challenging evaluation benchmarks to assess their ability to understand complex scientific data. Scientific tables, as core carriers of knowledge representation, combine text, symbols, and graphics, forming a typical multimodal reasoning scenario. However, existing benchmarks are mostly focused on general domains, failing to reflect the unique structural complexity and domain-specific semantics inherent in scientific research. Chemical tables are particularly representative: they intertwine structured variables such as reagents, conditions, and yields with visual symbols like molecular structures and chemical formulas, posing significant challenges to models in cross-modal alignment and semantic parsing. To address this, we propose ChemTable-a large scale benchmark of chemical tables constructed from real-world literature, containing expert-annotated cell layouts, logical structures, and domain-specific labels. It supports two core tasks: (1) table recognition (structure and content extraction); and (2) table understanding (descriptive and reasoning-based question answering). Evaluation on ChemTable shows that while mainstream multimodal models perform reasonably well in layout parsing, they still face significant limitations when handling critical elements such as molecular structures and symbolic conventions. Closed-source models lead overall but still fall short of human-level performance. This work provides a realistic testing platform for evaluating scientific multimodal understanding, revealing the current bottlenecks in domain-specific reasoning and advancing the development of intelligent systems for scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。