用智能流程自动提取文献中的蛋白降解数据,提升数据库规模与质量。
Beyond Manual Curation: Augmenting Targeted Protein Degradation Databases via Agentic Literature Extraction Workflows

- 构建专家协作的LLM工作流,精准提取化合物、靶点、实验条件等关键信息。
- 仅用7篇标注文献,即实现98%准确率,扩展数据库超八成新记录。
- 适合药物研发、AI辅助科研人员,助力高效构建高质量生物数据集。
生物医学预测模型依赖于藏在论文正文、表格和补充材料中的结构化实验数据。这一瓶颈在靶向蛋白降解(TPD)领域尤为突出,因每条记录需整合化合物身份、降解靶点、招募子、实验背景及终点值,这些信息分散在不同部分。化合物命名不一致及实验背景不完整,使通用大模型难以胜任。现有分子胶和PROTAC数据库多为人工整理,常缺必要实验上下文。本文将TPD数据库提取定义为领域专用的梳理任务,提出一种专家参与的LLM工作流,并通过三重对比(LLM预测、标准基线、专家标注真值)评估。轻量级交叉验证提示优化模块,仅凭少量专家标注即可调整提取指令。仅用7篇分子胶文献标注,便达到记录级F1=0.98;通过术语替换迁移至PROTAC,仍保持记录级F1>0.93。规模化应用后,分子胶与PROTAC数据库分别扩充81%和92%记录,其中92%与82.5%的新记录经专家验证正确。该工作流还恢复了动力学与实验上下文信息,支持跨研究效力比较与条件感知建模。相关工作流、提示模板、评估代码及提取数据集均已开源,可供更广泛科学数据整理与AI辅助研究使用。
原文摘要 · Abstract (English)
Predictive models in biomedicine depend on structured assay data locked in the text, tables, and supplements of primary publications. This bottleneck is especially acute in targeted protein degradation (TPD), where each assay record must combine compound identity, degradation target, recruiter, assay context, and endpoint values reported across sections, tables, and supplementary files. Inconsistent compound identifiers and incomplete or implicit assay context further demand domain-specific logic that generic LLM pipelines do not provide. Existing molecular glue and PROTAC databases are manually curated and often lack the experimental context required for downstream modeling. We formulate TPD database extraction as a domain-specific curation task and present an expert-in-the-loop LLM workflow, evaluated through a triangular comparison among LLM predictions, standardized baseline records, and expert-annotated ground truth. A lightweight cross-validated prompt-refinement module adapts extraction instructions from scarce expert annotations. With only seven annotated molecular glue publications, the workflow achieved record-level $F_1 = 0.98$ and transferred to PROTACs by terminology substitution alone, maintaining record-level $F_1 > 0.93$. Applied at scale, it expanded molecular glue and PROTAC databases by 81% and 92% records, respectively, with 92% and 82.5% of newly recovered records validated as correct upon expert review. The workflow also recovered kinetic and assay-context information essential for cross-study potency comparison and condition-aware degradation modeling. We release the workflow, prompts, evaluation code, and extracted datasets as resources for TPD data curation and AI-assisted scientific curation more broadly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。