arXiv:2501.15415cs.CV2025-01被引 9

将分子结构图转化为可读字符串,助力化学发现

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery

  • 提出OCSU任务,从图像理解扩展到多层级语义生成
  • 构建首个大规模数据集Vis-CheBI20,支持端到端模型训练
  • 双路径方案兼顾精度与泛化,适合化学与AI交叉研究者

从分子图形表示中理解化学结构是一项极具挑战性的图像描述任务,对分子中心的科学发现具有重要意义。分子图像和描述子任务的多样性给图像表征学习与任务建模带来显著困难。现有方法仅关注将分子图像转换为图结构的特定任务(即OCSR)。本文提出光学化学结构理解(OCSU)任务,将低层识别拓展至多层级理解,旨在将化学结构图转化为机器与化学家均可读的字符串。为推动OCSU技术发展,我们探索基于OCSR与无OCSR两种范式。提出DoubleCheck模型,通过注意力特征增强解决局部原子模糊问题,可与现有基于SMILES的方法级联实现OCSU。同时设计了端到端优化的Mol-VL视觉语言模型。构建了首个大规模OCSU数据集Vis-CheBI20。大量实验证明,所提方法能有效生成化学家可读的分子图描述,为后续研究提供坚实基线。代码、模型与数据已开源。

原文摘要 · Abstract (English)

Understanding the chemical structure from a graphical representation of a molecule is a challenging image caption task that would greatly benefit molecule-centric scientific discovery. Variations in molecular images and caption subtasks pose a significant challenge in both image representation learning and task modeling. Yet, existing methods only focus on a specific caption task that translates a molecular image into its graph structure, i.e., OCSR. In this paper, we propose the Optical Chemical Structure Understanding (OCSU) task, which extends low-level recognition to multilevel understanding and aims to translate chemical structure diagrams into readable strings for both machine and chemist. To facilitate the development of OCSU technology, we explore both OCSR-based and OCSR-free paradigms. We propose DoubleCheck to enhance OCSR performance via attentive feature enhancement for local ambiguous atoms. It can be cascaded with existing SMILES-based molecule understanding methods to achieve OCSU. Meanwhile, Mol-VL is a vision-language model end-to-end optimized for OCSU. We also construct Vis-CheBI20, the first large-scale OCSU dataset. Through comprehensive experiments, we demonstrate the proposed approaches excel at providing chemist-readable caption for chemical structure diagrams, which provide solid baselines for further research. Our code, model, and data are open-sourced at https://github.com/PharMolix/OCSU.

化学图像理解分子生成视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。