arXiv:2604.06685cs.CLcs.AI2026-04ACL被引 2

让化学视觉模型先推理再回答,提升可解释性

ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding

  • 先识别官能团等化学细节,再生成答案,强化推理过程
  • 在76万样本数据上训练,性能超越主流商业与开源模型
  • 适合需要可解释化学分析的研究者和开发者

尽管视觉语言模型(VLMs)在化学视觉理解中展现巨大潜力,现有模型主要针对直接问答任务优化,导致系统缺乏可解释性。本文提出ChemVLR,一种优先在感知阶段进行推理的化学视觉语言模型。不同于传统方法,ChemVLR通过显式识别官能团等细粒度化学描述符来分析视觉输入,从而生成清晰可解释的推理路径。为此,我们采用跨模态逆向工程策略与严格筛选流程,构建了包含76万条高质量样本的大规模推理与标注数据集,涵盖分子与反应任务。同时,采用三阶段训练框架系统提升模型感知与推理能力。实验表明,ChemVLR达到当前最优(SOTA)性能,超越领先专有模型及领域特定开源基线。我们还进行了全面消融实验,验证训练策略与数据生成设计的有效性。代码与模型权重将发布于https://github.com/xxlllz/ChemVLR。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) have demonstrated significant potential in chemical visual understanding, current models are predominantly optimized for direct visual question-answering tasks. This paradigm often results in "black-box" systems that fail to utilize the inherent capability of Large Language Models (LLMs) to infer underlying reaction mechanisms. In this work, we introduce ChemVLR, a chemical VLM designed to prioritize reasoning within the perception process. Unlike conventional chemical VLMs, ChemVLR analyzes visual inputs in a fine-grained manner by explicitly identifying granular chemical descriptors, such as functional groups, prior to generating answers. This approach ensures the production of explicit and interpretable reasoning paths for complex visual chemical problems. To facilitate this methodology, we implement a cross-modality reverse-engineering strategy, combined with a rigorous filtering pipeline, to curate a large-scale reasoning-and-captioning dataset comprising 760k high-quality samples across molecular and reaction tasks. Furthermore, we adopt a three-stage training framework that systemically builds model perception and reasoning capacity. Experiments demonstrate that ChemVLR achieves state-of-the-art (SOTA) performance, surpassing both leading proprietary models and domain-specific open-source baselines. We also provide comprehensive ablation studies to validate our training strategy and data generation designs. Code and model weights will be available at https://github.com/xxlllz/ChemVLR.

化学视觉推理增强可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。