arXiv:2604.21890cs.CL2026-04

构建首个大规模人工标注的开放域事件抽取数据集,支持通用事件理解。

EVENT5Ws: A Large Dataset for Open-Domain Event Extraction from Documents

论文配图:EVENT5Ws: A Large Dataset for Open-Domain Event Extraction from Documents
图 1 · 摘自论文原文
  • 设计系统化标注流程,实现开放域事件要素的精准标注
  • 覆盖多样事件类型,模型在跨区域数据上表现良好
  • 适合从事事件抽取、大模型评估的研究者使用

事件抽取旨在从文本中识别事件的核心要素,对应急决策等任务至关重要。现有算法训练数据存在局限:封闭域数据事件类型覆盖不足,开放域缺乏大规模人工验证数据。为此,我们构建了EVENT5Ws——一个大规模、人工标注且经统计验证的开放域事件抽取数据集。通过系统化标注流程,我们提供了标注复杂性的实证分析。利用该数据集,我们评估了主流预训练大语言模型,并建立了新基准。结果表明,基于EVENT5Ws训练的模型可有效泛化至不同地理背景的数据集,展现了其在构建通用算法方面的潜力。最后,我们总结了数据集构建过程中的经验教训,为未来大规模数据集开发提供参考。

原文摘要 · Abstract (English)

Event extraction identifies the central aspects of events from text. It supports event understanding and analysis, which is crucial for tasks such as informed decision-making in emergencies. Therefore, it is necessary to develop automated event extraction approaches. However, existing datasets for algorithm development have limitations, including limited coverage of event types in closed-domain settings and a lack of large, manually verified dataset in open-domain settings. To address these limitations, we create EVENT5Ws , a large, manually annotated, and statistically verified open-domain event extraction dataset. We design a systematic annotation pipeline to create the dataset and provide empirical insights into annotation complexity. Using EVENT5Ws, we evaluate state-of-the-art pre-trained large language models and establish a benchmark for future research. We further show that models trained on EVENT5Ws generalize effectively to datasets from different geographical contexts, which demonstrates its potential for developing generalizable algorithms. Finally, we summarize the lessons learned during the dataset development and provide recommendations to support future large-scale dataset development.

事件抽取数据集开放域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。