端到端识别化学文献中的多模态结构,提升药物研发效率
MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
- 结合图文布局信息,用双阶段训练融合多模态编码
- 在真实数据集上准确率超越现有方法32.7个百分点
- 适合药物发现、化学文献挖掘等领域的研究人员
从文档中自动提取化学结构是大规模分析化学文献的关键。现有方法仅能独立处理图或文本中的分子结构,而针对多模态描述(马库什结构)的识别精度不足,难以用于大规模自动化处理。本文提出MarkushGrapher-2,一种端到端的多模态化学结构识别方法:首先使用专用OCR模型从化学图像中提取文本;其次通过视觉-文本-布局编码器与光学化学结构识别视觉编码器联合编码文本、图像和布局信息;最后采用两阶段训练策略融合编码结果,实现自回归生成马库什结构表示。为解决训练数据短缺问题,我们构建了大规模真实世界马库什结构数据集的自动构建流程,并发布IP5-M——一个大型人工标注的真实世界马库什结构基准数据集,以推动该任务研究。大量实验表明,本方法在多模态马库什结构识别上显著优于现有最先进模型,同时保持了优异的分子结构识别性能。代码、模型与数据集已公开。
原文摘要 · Abstract (English)
Automatically extracting chemical structures from documents is essential for the large-scale analysis of the literature in chemistry. Automatic pipelines have been developed to recognize molecules represented either in figures or in text independently. However, methods for recognizing chemical structures from multimodal descriptions (Markush structures) lag behind in precision and cannot be used for automatic large-scale processing. In this work, we present MarkushGrapher-2, an end-to-end approach for the multimodal recognition of chemical structures in documents. First, our method employs a dedicated OCR model to extract text from chemical images. Second, the text, image, and layout information are jointly encoded through a Vision-Text-Layout encoder and an Optical Chemical Structure Recognition vision encoder. Finally, the resulting encodings are effectively fused through a two-stage training strategy and used to auto-regressively generate a representation of the Markush structure. To address the lack of training data, we introduce an automatic pipeline for constructing a large-scale dataset of real-world Markush structures. In addition, we present IP5-M, a large manually-annotated benchmark of real-world Markush structures, designed to advance research on this challenging task. Extensive experiments show that our approach substantially outperforms state-of-the-art models in multimodal Markush structure recognition, while maintaining strong performance in molecule structure recognition. Code, models, and datasets are released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。