BioMiner自动从文献中提取蛋白质-配体活性数据,提升药物研发效率。
BioMiner: A Multi-modal System for Automated Mining of Protein-Ligand Bioactivity Data from Literature

- 分步处理:先理解生物活性语义,再用化学视觉推理构建精确分子结构。
- 在16,457条数据上实现0.32的三元组F1分数,验证了提取能力。
- 适合药物发现、化学信息学研究者,尤其关注自动化数据挖掘的团队。
发表于文献中的蛋白质-配体生物活性数据对药物发现至关重要,但人工整理难以跟上文献增长速度。自动化提取面临挑战,因需解析文本、表格和图表中分散的生化语义,并重建精确的配体结构(如马克什结构)。为解决这一瓶颈,我们提出BioMiner,一种多模态提取框架,将生物活性语义理解与配体结构构建分离处理。其中,生物活性语义通过直接推理获取,化学结构则基于化学引导的视觉语义推理范式,由多模态大模型在化学基础的视觉表征上推断结构间关系,并将精确分子构建交由领域化学工具完成。为严格评估与方法开发,我们建立了一个包含16,457个生物活性条目的基准BioVista,涵盖500篇文献。BioMiner验证了其提取能力并提供量化基线,三元组F1得分为0.32。实际应用显示:(1) 从11,683篇论文中提取82,262条数据,构建预训练数据库,使下游模型性能提升3.9%;(2) 支持人机协同工作流,使高质量NLRP3活性数据量翻倍,相较28个QSAR模型提升38.6%,识别出16个具有新骨架的候选化合物;(3) 加速蛋白-配体复合物活性标注,在PoseBusters数据集上实现5.59倍提速和5.75%准确率提升。
原文摘要 · Abstract (English)
Protein-ligand bioactivity data published in the literature are essential for drug discovery, yet manual curation struggles to keep pace with rapidly growing literature. Automated bioactivity extraction remains challenging because it requires not only interpreting biochemical semantics distributed across text, tables, and figures, but also reconstructing chemically exact ligand structures (e.g., Markush structures). To address this bottleneck, we introduce BioMiner, a multi-modal extraction framework that explicitly separates bioactivity semantic interpretation from ligand structure construction. Within BioMiner, bioactivity semantics are inferred through direct reasoning, while chemical structures are resolved via a chemical-structure-grounded visual semantic reasoning paradigm, in which multi-modal large language models operate on chemically grounded visual representations to infer inter-structure relationships, and exact molecular construction is delegated to domain chemistry tools. For rigorous evaluation and method development, we further establish BioVista, a comprehensive benchmark comprising 16,457 bioactivity entries curated from 500 publications. BioMiner validates its extraction ability and provides a quantitative baseline, achieving an F1 score of 0.32 for bioactivity triplets. BioMiner's practical utility is demonstrated via three applications: (1) extracting 82,262 data from 11,683 papers to build a pre-training database that improves downstream models performance by 3.9%; (2) enabling a human-in-the-loop workflow that doubles the number of high-quality NLRP3 bioactivity data, helping 38.6% improvement over 28 QSAR models and identification of 16 hit candidates with novel scaffolds; and (3) accelerating protein-ligand complex bioactivity annotation, achieving a 5.59-fold speed increase and 5.75% accuracy improvement over manual workflows in PoseBusters dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。