构建首个聚焦肠脑轴的精细生物医学关系抽取基准数据集
A Domain-Specific Curated Benchmark for Entity and Document-Level Relation Extraction
- 基于1600+篇文献,专家人工标注细粒度实体与概念关联
- 涵盖实体识别、链接及多层级关系抽取,支持跨领域评估
- 兼顾高质量标注与弱监督数据,适合训练和测试生物医学信息提取系统
信息抽取(IE)包括命名实体识别(NER)、命名实体链接(NEL)和关系抽取(RE),对将快速增长的科学出版物转化为结构化知识至关重要。这一需求在肠道-大脑轴等快速演进的生物医学领域尤为突出,相关研究探索肠道微生物群与脑部疾病间的复杂互动。然而,现有生物医学信息抽取基准往往范围狭窄,依赖远距离监督或自动生成标注,限制了其对鲁棒信息抽取方法发展的推动作用。我们提出了GutBrainIE,基于超过1600篇PubMed摘要,由生物医学与术语学专家手工标注细粒度实体、概念级链接和关系。尽管源自肠脑轴主题,该基准具备丰富的模式体系、多项任务设计以及高质量与弱监督数据的结合,可广泛应用于不同领域生物医学信息抽取系统的开发与评估。
原文摘要 · Abstract (English)
Information Extraction (IE), encompassing Named Entity Recognition (NER), Named Entity Linking (NEL), and Relation Extraction (RE), is critical for transforming the rapidly growing volume of scientific publications into structured, actionable knowledge. This need is especially evident in fast-evolving biomedical fields such as the gut-brain axis, where research investigates complex interactions between the gut microbiota and brain-related disorders. Existing biomedical IE benchmarks, however, are often narrow in scope and rely heavily on distantly supervised or automatically generated annotations, limiting their utility for advancing robust IE methods. We introduce GutBrainIE, a benchmark based on more than 1,600 PubMed abstracts, manually annotated by biomedical and terminological experts with fine-grained entities, concept-level links, and relations. While grounded in the gut-brain axis, the benchmark's rich schema, multiple tasks, and combination of highly curated and weakly supervised data make it broadly applicable to the development and evaluation of biomedical IE systems across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。