构建化学多模态推理评测基准,检验模型跨视觉、文本、符号融合能力。
ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry
- 设计三模态输入:图像、图文混合、SMILES符号,评估跨模态推理
- 视觉输入下模型表现差,结构化学最困难,多模态融合仍存错误
- 自动化评估流程可诊断失败原因,适合化学与AI交叉研究者使用
化学推理天然融合视觉、文本与符号模态,但现有评测基准常依赖简单图像-文本对,缺乏化学语义深度。为此,我们提出 extbf{ChemVTS-Bench},一个领域真实、系统化的多模态大模型(MLLM)评测基准,涵盖有机分子、无机材料与三维晶体结构等复杂化学问题。每项任务提供三种输入模式:(1) 仅视觉,(2) 视觉-文本混合,(3) 基于SMILES的符号输入,支持细粒度分析模态依赖行为与跨模态融合。我们开发自动化代理工作流,实现标准化推理、答案验证与故障诊断。大量实验表明:仅视觉输入仍具挑战性,结构化学最难,多模态融合虽缓解视觉、知识或逻辑错误,但未根除。该基准为推进多模态化学推理提供了严谨、领域忠实的测试平台。所有数据与代码将开源。
原文摘要 · Abstract (English)
Chemical reasoning inherently integrates visual, textual, and symbolic modalities, yet existing benchmarks rarely capture this complexity, often relying on simple image-text pairs with limited chemical semantics. As a result, the actual ability of Multimodal Large Language Models (MLLMs) to process and integrate chemically meaningful information across modalities remains unclear. We introduce \textbf{ChemVTS-Bench}, a domain-authentic benchmark designed to systematically evaluate the Visual-Textual-Symbolic (VTS) reasoning abilities of MLLMs. ChemVTS-Bench contains diverse and challenging chemical problems spanning organic molecules, inorganic materials, and 3D crystal structures, with each task presented in three complementary input modes: (1) visual-only, (2) visual-text hybrid, and (3) SMILES-based symbolic input. This design enables fine-grained analysis of modality-dependent reasoning behaviors and cross-modal integration. To ensure rigorous and reproducible evaluation, we further develop an automated agent-based workflow that standardizes inference, verifies answers, and diagnoses failure modes. Extensive experiments on state-of-the-art MLLMs reveal that visual-only inputs remain challenging, structural chemistry is the hardest domain, and multimodal fusion mitigates but does not eliminate visual, knowledge-based, or logical errors, highlighting ChemVTS-Bench as a rigorous, domain-faithful testbed for advancing multimodal chemical reasoning. All data and code will be released to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。