构建可查询的实验室工作流知识图谱,捕捉专家隐性经验
Federated Semantic Knowledge Graphs for Laboratory Workflows: A Structured Expert Elicitation Methodology Demonstrated Through Bioanalytical Workflow Twins

- 用AI访谈代理结构化提取专家判断,生成三层知识图谱
- 跨子图查询发现自动化系统掩盖的隐性失败,提升科学有效性
- 为实验室AI agent提供可推理的失败模式语义模型,适合制药研发
生物医药研究中的实验流程包含大量隐性知识——如失败条件判断、决策分支逻辑和上下文依赖关系——这些信息无法通过协议文档、传感器数据或现有生物医学本体获取。本文提出一种可重复的结构化专家萃取方法与联邦语义知识图谱(SKG)架构,在基因泰克生化与细胞药理部门完成部署。通过专设的协议智能协作者AI访谈代理,应用结构化萃取视角,以专家赋信度评分的方式挖掘程序性知识,生成涵盖项目级决策节点、检测协议知识及物理执行基础设施的三层图谱。分别构建的子图(如免疫分析(ELISA)、定量质谱(LC-MS/PRM)、实验室自动化)通过共享上层本体对齐,并作为单一联邦图谱进行查询。评估显示,该系统支持七类原有数据源无法实现的查询,包括跨子图遍历,可识别自动化系统掩盖的隐性失败——即执行日志显示成功但科学有效性受损的情况。关键的MASKED_BY关系编码了当前信息平台无法感知的一类实验室风险:即现有系统无法推理科学有效性的结构性缺陷。该架构提供了当前实验室AI代理所缺乏的语义世界模型,可查询工作流中静默失效的位置、人类判断不可替代的环节,以及哪些执行资产会掩盖而非检测失败。
原文摘要 · Abstract (English)
Laboratory workflows in pharmaceutical and biomedical research encode substantial tacit knowledge -- expert judgment about failure conditions, decision branching logic, and contextual dependencies -- that remains inaccessible to protocol documents, sensor streams, and existing biomedical ontologies. We present a repeatable structured expert elicitation methodology and federated Semantic Knowledge Graph (SKG) architecture for capturing and querying this knowledge, demonstrated through deployment at the Biochemical and Cellular Pharmacology Department of Genentech. Knowledge is elicited via the Protocol Intelligence Co-pilot, a purpose-built AI interview agent that applies structured elicitation lenses to surface tacit procedural knowledge with expert-assigned confidence scores, producing graph representations across three tiers: program-level decision milestones, assay protocol knowledge, and physical execution infrastructure. Separately constructed subgraphs, exemplified by immunoassay (ELISA), quantitative mass spectrometry (LC-MS/PRM), and laboratory automation, are aligned through a shared upper ontology and queried as a single federated graph. Evaluation demonstrates seven query types structurally unavailable from any individual data source, including a cross-subgraph traversal that identifies automation-masked silent failures -- conditions where execution logs report success while scientific validity is compromised. Critically, the MASKED_BY graph relationship encodes a class of laboratory risk invisible to current informatics platforms -- the structural gap that prevents existing systems from reasoning about scientific validity. This architecture provides the semantic world model that AI laboratory agents currently lack: a queryable representation of where workflows fail silently, where human judgment is irreplaceable, and which execution assets mask rather than detect failure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。