提升化学结构图像转文本的准确率,尤其改进了复杂结构的识别。
MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition

- 将化学结构识别视为图像到文本的翻译任务,统一处理单分子与马库什结构。
- 提出OCSRGlyph模型,在立体化学信息上优化,显著提升单分子识别精度。
- 首次引入新评估指标,更准确衡量马库什结构翻译的正确性,适合专利分析者。
化学结构常以图像形式出现在专利和科学文献中。为实现程序化使用(如数据库索引或机器学习训练集构建),需将其转换为线性表示。该任务主要有两类:单个分子的光学化学结构识别(OCSR)和代表一类分子的马库什结构解析。尽管前者已较为成熟,马库什结构解析仍是难点。本文将两类任务均视为图像到文本的翻译问题。提出OCSRGlyph,一种先进的OCSR模型,在考虑立体化学的基础上提升了性能。针对马库什结构,提出MarkushGlyph,一个视觉-语言模型,可一次性读取整个马库什结构图像,区别于以往分步处理视觉与文本内容的方法。最后,设计了一种新评估指标,有效处理先前指标存在的误判问题。
原文摘要 · Abstract (English)
Chemical structures appear in patents and the scientific literature as images. For programmatic usage, such as indexing in databases or constructing machine learning model training sets, they must be transformed into line notations. The two common forms of this task are translating an image of a single molecule (optical chemical structure recognition - OCSR) and translating a Markush structure that represents a family of molecules. While prior work in the former case is quite mature, Markush structure parsing remains a challenging task. In this work, we treat both tasks as an image-to-text translation problem. We then propose OCSRGlyph, a state-of-the-art OCSR model, improving performance over prior methods by carefully considering stereochemistry. For the Markush task, we introduce MarkushGlyph, a vision-language model that reads the entire Markush structure as an image. This contrasts with prior systems, which often use multiple stages to separately process visual and text input content. Finally, we introduce a new metric for determining the accuracy of Markush structure translations, handling failure modes present in prior metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。