一键提取文献中的分子结构与反应信息,自动化程度高
MolMole: Molecule Mining from Scientific Literature
- 基于视觉的统一框架,端到端处理分子检测与反应图解析
- 在550页标注数据集上表现优于现有工具,准确率显著提升
- 适合化学信息学、药物研发人员快速获取文献化学数据
从科学文献中提取分子结构和反应数据面临格式多样、文档布局复杂等挑战。为此,我们提出MolMole,一个基于视觉的深度学习框架,将分子检测、反应图解析与光学化学结构识别(OCSR)整合为单一流程,实现从页面级文档直接自动化提取化学数据。针对缺乏标准页面级基准和评估指标的问题,我们构建了包含550页的测试集,每页标注分子边界框、反应标签和MOLfiles,并提出一种新型评估指标。实验表明,MolMole在自建基准及公开数据集上均优于现有工具包。测试集将公开,MolMole工具包也将通过LG AI Research官网提供交互式演示。商业合作请联系[email protected]。
原文摘要 · Abstract (English)
The extraction of molecular structures and reaction data from scientific documents is challenging due to their varied, unstructured chemical formats and complex document layouts. To address this, we introduce MolMole, a vision-based deep learning framework that unifies molecule detection, reaction diagram parsing, and optical chemical structure recognition (OCSR) into a single pipeline for automating the extraction of chemical data directly from page-level documents. Recognizing the lack of a standard page-level benchmark and evaluation metric, we also present a testset of 550 pages annotated with molecule bounding boxes, reaction labels, and MOLfiles, along with a novel evaluation metric. Experimental results demonstrate that MolMole outperforms existing toolkits on both our benchmark and public datasets. The benchmark testset will be publicly available, and the MolMole toolkit will be accessible soon through an interactive demo on the LG AI Research website. For commercial inquiries, please contact us at \href{mailto:[email protected]}{contact\[email protected]}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。