构建多语言冲突事件数据集,提升跨语言事件理解能力。
LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World
- 提出抽象事件抽取与实体链接新方法,基于文档整体理解
- 零样本系统在端到端任务上达58.3% F1,实体链接达45.7% F1
- 适用于全球冲突分析、多语言信息整合的研究者
本文提出 LEMONADE,一个涵盖20种语言、171个国家的大型冲突事件数据集,共包含39,786个事件,覆盖大量区域性实体。该数据集基于部分重标注的武装冲突地点与事件数据(ACLED)构建。为应对多语言来源聚合难题,我们引入抽象事件抽取(AEE)及其子任务抽象实体链接(AEL)。不同于传统的基于片段的事件抽取,本方法通过全局文档理解检测事件要素并实现跨语言归一化。我们在多种大语言模型(LLMs)上评估了这些任务,适配现有零样本事件抽取系统,并基准化监督模型。此外,我们提出 ZEST,一种新型零样本检索式 AEL 系统。最佳零样本系统在端到端任务中取得58.3% F1,LLMs 超越 GoLLIE 等专用模型;在实体链接任务中,ZEST 达到45.7% F1,显著优于仅有23.7% F1的 OneNet 基线。然而,零样本结果仍分别落后于最优监督系统20.1%和37.0%,凸显进一步研究必要性。
原文摘要 · Abstract (English)
This paper presents LEMONADE, a large-scale conflict event dataset comprising 39,786 events across 20 languages and 171 countries, with extensive coverage of region-specific entities. LEMONADE is based on a partially reannotated subset of the Armed Conflict Location & Event Data (ACLED), which has documented global conflict events for over a decade. To address the challenge of aggregating multilingual sources for global event analysis, we introduce abstractive event extraction (AEE) and its subtask, abstractive entity linking (AEL). Unlike conventional span-based event extraction, our approach detects event arguments and entities through holistic document understanding and normalizes them across the multilingual dataset. We evaluate various large language models (LLMs) on these tasks, adapt existing zero-shot event extraction systems, and benchmark supervised models. Additionally, we introduce ZEST, a novel zero-shot retrieval-based system for AEL. Our best zero-shot system achieves an end-to-end F1 score of 58.3%, with LLMs outperforming specialized event extraction models such as GoLLIE. For entity linking, ZEST achieves an F1 score of 45.7%, significantly surpassing OneNet, a state-of-the-art zero-shot baseline that achieves only 23.7%. However, these zero-shot results lag behind the best supervised systems by 20.1% and 37.0% in the end-to-end and AEL tasks, respectively, highlighting the need for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。