arXiv:2601.04524cs.AI2026-01中稿 · EMNLP

构建首个面向实验理解的生物医学流程知识图谱数据集。

BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding

  • 基于实验流程构建细粒度知识图谱,捕捉实体、动作与关系。
  • 在高信息密度和多步推理任务上,模型性能显著提升。
  • 适合需要精细实验理解的研究者与实验室自动化系统开发人员。

理解生物医学实验是下游任务(如实验室自动化)的基础,有助于跨学科沟通。但高信息密度(HID)和多步推理(MSR)带来挑战。现有生物医学结构化知识抽取数据集多为粗粒度,难以支持细粒度实验理解。为此,我们提出生物医学流程信息抽取数据集(BioPIE),提供以流程为中心的知识图谱(KG),涵盖足够规模的实体、动作与关系,支持跨实验协议的推理。我们在BioPIE上评估了监督学习与大模型(LLM)方法,验证其有效性,并构建了一个生物医学问答系统,定量展示其在下游理解任务中的优势。实验结果表明,模型在HID与MSR问题集上的理解性能均有提升。

原文摘要 · Abstract (English)

Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cross-disciplinary communication. Two challenges, High Information Density (HID) and Multi-Step Reasoning (MSR), pose unique difficulties for precise automatic experimental understanding. Extracting structured knowledge, e.g., Knowledge Graphs (KGs), is an effective approach to address the HID and MSR. However, existing biomedical datasets for structured knowledge Information Extraction (IE) are limited to a general or coarse-grained level, hindering fine-grained experimental understanding. To address this gap, we introduce Biomedical Protocol Information Extraction Dataset (BioPIE), a dataset providing procedure-centric KGs that captures entities, actions, and relations at a scale sufficient for reasoning across biomedical protocols. We evaluate both supervised and LLM-based IE methods on BioPIE to verify its effectiveness, and implement a biomedical question answering system to provide a quantitative illustration of BioPIE's effectiveness for downstream understanding tasks. The experimental results demonstrate improved understanding performance on both the HID and MSR question sets.

知识图谱实验理解生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。