自动从文献中提取酶反应数据,助力人工智能驱动的生物化学研究
zERExtractor:An Automated Platform for Enzyme-Catalyzed Reaction Data Extraction from Scientific Literature
- 融合深度学习与大模型,模块化设计支持持续升级
- 在表格识别、分子图像解析等任务上准确率超90%
- 适合从事酶动力学建模与知识发现的研究人员
酶动力学文献的快速增长已超出主流生化数据库的人工整理能力,成为AI驱动建模与知识发现的主要障碍。我们提出zERExtractor,一个自动化且可扩展的平台,用于从科学文献中全面提取酶催化反应与活性数据。该平台采用统一模块化架构,支持先进模型(包括大语言模型)的即插即用,实现系统随AI进展持续演进。其流程结合领域适配的深度学习、高级OCR、语义实体识别和提示驱动的LLM模块,并引入专家校正,可从异构文档中提取动力学参数(如kcat、Km)、酶序列、底物SMILES、实验条件及分子图。通过集成AI辅助标注、专家验证与迭代优化的主动学习策略,系统能快速适应新数据源。我们还发布了涵盖270篇关于P450酶学文献的大规模基准数据集,包含超过1,000个标注表格和5,000个生物字段。基准测试显示,zERExtractor在表格识别(准确率89.9%)、分子图像解析(最高达99.1%)和关系抽取(准确率94.2%)方面均优于现有基线。该平台为酶动力学数据鸿沟提供灵活、高保真提取框架,为未来基于AI的酶建模与生化知识发现奠定基础。
原文摘要 · Abstract (English)
The rapid expansion of enzyme kinetics literature has outpaced the curation capabilities of major biochemical databases, creating a substantial barrier to AI-driven modeling and knowledge discovery. We present zERExtractor, an automated and extensible platform for comprehensive extraction of enzyme-catalyzed reaction and activity data from scientific literature. zERExtractor features a unified, modular architecture that supports plug-and-play integration of state-of-the-art models, including large language models (LLMs), as interchangeable components, enabling continuous system evolution alongside advances in AI. Our pipeline combines domain-adapted deep learning, advanced OCR, semantic entity recognition, and prompt-driven LLM modules, together with human expert corrections, to extract kinetic parameters (e.g., kcat, Km), enzyme sequences, substrate SMILES, experimental conditions, and molecular diagrams from heterogeneous document formats. Through active learning strategies integrating AI-assisted annotation, expert validation, and iterative refinement, the system adapts rapidly to new data sources. We also release a large benchmark dataset comprising over 1,000 annotated tables and 5,000 biological fields from 270 P450-related enzymology publications. Benchmarking demonstrates that zERExtractor consistently outperforms existing baselines in table recognition (Acc 89.9%), molecular image interpretation (up to 99.1%), and relation extraction (accuracy 94.2%). zERExtractor bridges the longstanding data gap in enzyme kinetics with a flexible, plugin-ready framework and high-fidelity extraction, laying the groundwork for future AI-powered enzyme modeling and biochemical knowledge discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。