构建跨领域科学事件抽取基准,提升科研文本理解能力。
SciEvent: Benchmarking Multi-domain Scientific Event Extraction
- 统一事件抽取框架,分两阶段提取科学活动与关键要素。
- 涵盖5个领域500篇论文,标注事件片段、触发词与细粒度论元。
- 暴露现有模型在社科人文领域的不足,推动多领域通用化研究。
科学信息抽取(SciIE)长期依赖窄领域实体关系抽取,难以适用于跨学科研究,常因上下文缺失导致信息碎片化或矛盾。本文提出SciEvent,一个基于统一事件抽取(EE)范式的多领域科学摘要基准,包含5个研究领域共500篇摘要,人工标注了事件片段、触发词及细粒度论元。我们将SciIE定义为多阶段流程:(1)将摘要分割为背景、方法、结果、结论四类核心科学活动;(2)提取对应触发词与论元。通过微调的事件抽取模型、大语言模型(LLMs)及人工标注者的实验发现,当前模型在社会学和人文学科等领域的表现存在显著差距。SciEvent作为挑战性基准,推动可泛化的多领域科学信息抽取发展。
原文摘要 · Abstract (English)
Scientific information extraction (SciIE) has primarily relied on entity-relation extraction in narrow domains, limiting its applicability to interdisciplinary research and struggling to capture the necessary context of scientific information, often resulting in fragmented or conflicting statements. In this paper, we introduce SciEvent, a novel multi-domain benchmark of scientific abstracts annotated via a unified event extraction (EE) schema designed to enable structured and context-aware understanding of scientific content. It includes 500 abstracts across five research domains, with manual annotations of event segments, triggers, and fine-grained arguments. We define SciIE as a multi-stage EE pipeline: (1) segmenting abstracts into core scientific activities--Background, Method, Result, and Conclusion; and (2) extracting the corresponding triggers and arguments. Experiments with fine-tuned EE models, large language models (LLMs), and human annotators reveal a performance gap, with current models struggling in domains such as sociology and humanities. SciEvent serves as a challenging benchmark and a step toward generalizable, multi-domain SciIE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。