arXiv:2605.17214cs.AIcs.CL2026-05

让大模型读懂化学反应图,准确率提升20个百分点。

ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding

论文配图:ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding
图 1 · 摘自论文原文
  • 用视觉锚点定位官能团,结合语义对齐激活模型化学推理能力。
  • 在OCRDBench上实现92.0%的结构识别准确率,性能平均提升20个百分点。
  • 适合需要复杂化学推理的开放模型研究者与工业应用开发者。

尽管大语言模型(LLMs)已革新科学文本处理,但在解读化学反应图时仍存在显著能力差距。我们识别出两大瓶颈:视觉缺陷——通用视觉编码器难以解析密集分子图的严格拓扑连接;语义断层——标准线性字符串(如SMILES)无法有效激活模型的潜在化学推理能力。为此,我们提出化学视觉激活(ChemVA)框架,采用视觉锚点机制通过混合粒度检测定位官能团,并通过语义对齐将视觉特征转化为实体名称,以最大化激活LLMs中的知识。我们在新构建的OCRD-Bench数据集上评估该方法,该数据集包含密集的视觉-语义上下文和全面的反应覆盖范围,用于评估从识别到推理的全流程。大量实验表明,ChemVA在OCRD-Bench上达到92.0%的结构识别准确率,跨9种不同LLMs实现约20个百分点的持续性能提升,使开源模型在复杂化学推理任务中可媲美专有最先进系统。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have revolutionized scientific text processing, they exhibit a significant capability gap when interpreting chemical reaction diagrams. We identify two fundamental bottlenecks restricting current systems: a Visual Deficit, where generic vision encoders struggle to resolve the strict topological connectivity of dense molecular graphs, and a Semantic Disconnect, where standard linear strings, such as SMILES, fail to effectively activate the model's latent chemical reasoning. To bridge these gaps, we propose the Chemical Visual Activation (ChemVA) framework, which employs a Visual Anchor mechanism to ground functional groups via hybrid-granularity detection, followed by a semantic alignment approach that translates visual features into entity names to maximize knowledge activation in LLMs. We evaluate our approach on OCRD-Bench, a newly constructed dataset featuring dense visual-semantic contexts and comprehensive reaction coverage to evaluate the full spectrum from recognition to reasoning. Extensive experiments on OCRD-Bench demonstrate that ChemVA achieves 92.0% structural recognition accuracy. By bridging visual and semantic bottlenecks, our framework delivers a consistent performance gain of approximately 20 percentage points across 9 diverse LLMs, enabling open-weight models to rival proprietary SOTA systems in complex chemical reasoning tasks.

化学推理视觉-语言大模型图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。