arXiv:2411.11098cs.CV2024-11ICCV被引 21

端到端识别真实文献中的分子结构,尤其擅长处理复杂标记结构。

MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild

  • 用扩展SMILES规则标注数据,构建大规模分子图像数据集MolParser-7M。
  • 结合合成数据与真实专利文献样本,通过主动学习提升模型泛化能力。
  • 在复杂标记结构识别上超越现有方法,适合化学、药学领域研究者使用。

近年来,化学出版物和专利数量迅速增长,大量关键信息嵌入分子结构图中,给大规模文献检索带来挑战,并限制了大语言模型在生物、化学及制药领域的应用。精确提取化学结构至关重要。然而,现实文档中存在大量马克什结构,加之图像质量、绘图风格和噪声的差异,严重制约了现有光学化学结构识别(OCSR)方法的表现。我们提出MolParser,一种新型端到端OCSR方法,可高效准确地从真实文档中识别化学结构,包括复杂的马克什结构。我们采用扩展SMILES编码规则标注训练数据集,并基于此构建了目前最大规模的标注分子图像数据集MolParser-7M。在利用大量合成数据的同时,通过主动学习引入来自真实专利和科学文献的裁剪样本进行训练。我们采用课程学习方法训练端到端分子图像描述模型MolParser。该模型在多数场景下显著优于传统及学习型方法,具有广泛下游应用潜力。数据集已公开于HuggingFace。

原文摘要 · Abstract (English)

In recent decades, chemistry publications and patents have increased rapidly. A significant portion of key information is embedded in molecular structure figures, complicating large-scale literature searches and limiting the application of large language models in fields such as biology, chemistry, and pharmaceuticals. The automatic extraction of precise chemical structures is of critical importance. However, the presence of numerous Markush structures in real-world documents, along with variations in molecular image quality, drawing styles, and noise, significantly limits the performance of existing optical chemical structure recognition (OCSR) methods. We present MolParser, a novel end-to-end OCSR method that efficiently and accurately recognizes chemical structures from real-world documents, including difficult Markush structure. We use a extended SMILES encoding rule to annotate our training dataset. Under this rule, we build MolParser-7M, the largest annotated molecular image dataset to our knowledge. While utilizing a large amount of synthetic data, we employed active learning methods to incorporate substantial in-the-wild data, specifically samples cropped from real patents and scientific literature, into the training process. We trained an end-to-end molecular image captioning model, MolParser, using a curriculum learning approach. MolParser significantly outperforms classical and learning-based methods across most scenarios, with potential for broader downstream applications. The dataset is publicly available in huggingface.

分子识别图像理解化学信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。