自动化提取专利中的化学结构与活性数据,加速药物发现。
BioChemInsight: An Online Platform for Automated Extraction of Chemical Structures and Activity Data from Patents
- 整合多模型实现结构、活性、标识符三类数据自动抽取
- 在181篇专利上平均准确率超90%,支持15个靶点分析
- 适合药物研发、数据挖掘人员快速获取专利中被忽视的化学信息
自动化提取化学结构及其生物活性数据对加速药物发现和推动数据驱动研究至关重要。现有光学化学结构识别工具无法自主关联分子结构与活性谱,严重制约构效关系分析。为此,我们提出 BioChemInsight,一个开源流水线,集成 DECIMER Segmentation 与 MolNexTR 进行化学结构识别,使用 GLM-4.5V 关联化合物标识符,结合 PaddleOCR 与 GLM-4.6 完成生物活性数据提取与单位归一化。在覆盖15个治疗靶点的181篇专利上评估,系统在化学结构识别、活性数据提取和标识符关联三项任务中平均准确率均超过90%。分析表明,专利所覆盖的化学空间与公共数据库 ChEMBL 大部分互补。通过系统性专利挖掘,BioChemInsight 可获取在 ChEMBL 中未充分涵盖的化学信息,拓展可探索的化合物-靶点互作范围,丰富定量构效关系建模与靶向筛选的数据基础,并将数据预处理时间从数周缩短至数小时。项目代码已开源:https://github.com/dahuilangda/BioChemInsight。
原文摘要 · Abstract (English)
The automated extraction of chemical structures and their corresponding bioactivity data is essential for accelerating drug discovery and enabling data-driven research. Current optical chemical structure recognition tools lack the capability to autonomously link molecular structures with their bioactivity profiles, posing a significant bottleneck in structure-activity relationship analysis. To address this, we present BioChemInsight, an open-source pipeline that integrates DECIMER Segmentation with MolNexTR for chemical structure recognition, GLM-4.5V for compound identifier association, and PaddleOCR combined with GLM-4.6 for bioactivity extraction and unit normalization. We evaluated BioChemInsight on 181 patents covering 15 therapeutic targets. The system achieved an average extraction accuracy of above 90% across three key tasks: chemical structure recognition, bioactivity data extraction, and compound identifier association. Our analysis indicates that the chemical space covered by patents is largely complementary to that contained in established public database ChEMBL. Consequently, by enabling systematic patent mining, BioChemInsight provides access to chemical information underrepresented in ChEMBL. This capability expands the landscape of explorable compound-target interactions, enriches the data foundation for quantitative structure-activity relationship modeling and targeted screening, and reduces data preprocessing time from weeks to hours. BioChemInsight is available at https://github.com/dahuilangda/BioChemInsight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。