arXiv:2603.15011cs.CV2026-03被引 1

用分子编号做视觉提示,提升化学反应图解析准确率

Molecular Identifier Visual Prompt and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing

  • 用分子编号作为视觉提示,激活预训练模型的化学知识
  • 新强化学习方法在反应级指标上超越传统微调,提升显著
  • 适用于需要高精度解析化学文献的研究者和开发者

反应图解析(RxnDP)对从文献中提取化学合成信息至关重要。尽管视觉-语言模型(VLMs)为自动化这一复杂视觉推理任务提供了前景,但其应用受限于视觉化学实体与预训练知识无法对齐,以及词元级训练与反应级评估之间的固有差异。本文从提示表示和学习范式两个互补角度改进基于VLM的RxnDP:首先提出标识符视觉提示(IdtVP),利用自然出现的分子标识符(如加粗数字1a)激活VLM预训练中获取的化学知识,实现强大的零样本和分布外能力,优于现有提示策略;其次,提出Re3-DAPO强化学习算法,通过可验证奖励直接优化反应级指标,在微调范式下实现持续性能提升。此外,我们发布ScannedRxn基准数据集,包含带真实世界伪影的扫描历史反应图,用于严格评估模型鲁棒性和分布外能力。我们的工作推动了基于VLM的反应图解析在准确率和泛化性上的进步。数据、模型与代码将开源。

原文摘要 · Abstract (English)

Reaction diagram parsing (RxnDP) is critical for extracting chemical synthesis information from literature. Although recent Vision-Language Models (VLMs) have emerged as a promising paradigm to automate this complex visual reasoning task, their application is fundamentally bottlenecked by the inability to align visual chemical entities with pre-trained knowledge, alongside the inherent discrepancy between token-level training and reaction-level evaluation. To address these dual challenges, this work enhances VLM-based RxnDP from two complementary perspectives: prompting representation and learning paradigms. First, we propose Identifier as Visual Prompting (IdtVP), which leverages naturally occurring molecule identifiers (e.g., bold numerals like 1a) to activate the chemical knowledge acquired during VLM pre-training. IdtVP enables powerful zero-shot and out-of-distribution capabilities, outperforming existing prompting strategies. Second, to further optimize performance within fine-tuning paradigms, we introduce Re3-DAPO, a reinforcement learning algorithm that leverages verifiable rewards to directly optimize reaction-level metrics, thereby achieving consistent gains over standard supervised fine-tuning. Additionally, we release the ScannedRxn benchmark, comprising scanned historical reaction diagrams with real-world artifacts, to rigorously assess model robustness and out-of-distribution ability. Our contributions advance the accuracy and generalization of VLM-based reaction diagram parsing. We will release data, models, and code on GitHub.

化学信息学视觉语言模型强化学习图解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。