arXiv:2503.16096cs.CV2025-03CVPR被引 8

用图文布局联合识别专利中的复杂化学结构模板

MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures

  • 融合文本、图像和版面信息的多模态编码方法
  • 在真实数据集上达到领先性能,优于现有模型
  • 提供首个真实标注数据集,适合化学专利分析研究者

自动化分析化学文献有望加速材料科学与药物研发进程。尤其在专利文档中对化学结构及马库什结构(化学结构模板)的检索能力具有重要价值,如用于现有技术查新。尽管已有研究实现了从文本和图像中自动提取化学结构,但因马库什结构具有复杂的多模态特性,相关研究仍处于空白。本文提出 MarkushGrapher,一种针对文档中马库什结构的多模态识别方法。该方法通过视觉-文本-版面编码器与光学化学结构识别视觉编码器,联合编码文本、图像和布局信息,并自回归生成马库什结构的序列图表示及其变量组表格。为应对真实训练数据匮乏问题,我们设计了合成数据生成管道,可生成大量逼真的马库什结构。此外,我们构建了 M2S——首个真实世界马库什结构标注基准数据集,推动该难题的研究进展。大量实验表明,本方法在多数评估场景下优于现有的专用化学与通用视觉语言模型。代码、模型与数据集将公开。

原文摘要 · Abstract (English)

The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for chemical structures and Markush structures (chemical structure templates) within patent documents are valuable, e.g., for prior-art search. Advancements have been made in the automatic extraction of chemical structures from text and images, yet the Markush structures remain largely unexplored due to their complex multi-modal nature. In this work, we present MarkushGrapher, a multi-modal approach for recognizing Markush structures in documents. Our method jointly encodes text, image, and layout information through a Vision-Text-Layout encoder and an Optical Chemical Structure Recognition vision encoder. These representations are merged and used to auto-regressively generate a sequential graph representation of the Markush structure along with a table defining its variable groups. To overcome the lack of real-world training data, we propose a synthetic data generation pipeline that produces a wide range of realistic Markush structures. Additionally, we present M2S, the first annotated benchmark of real-world Markush structures, to advance research on this challenging task. Extensive experiments demonstrate that our approach outperforms state-of-the-art chemistry-specific and general-purpose vision-language models in most evaluation settings. Code, models, and datasets will be available.

化学结构多模态专利分析生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。