构建首个大规模中文事件中心型多文档摘要数据集,助力事件动态理解。
EventSum: A Large-Scale Event-Centric Summarization Dataset for Chinese Multi-News Documents
- 以百度百科为源,人工标注5100个事件的多新闻摘要任务
- 每事件平均11.4篇文档,总文本量达13,471字符,含复杂事件推理
- 设计事件、因果、时间等召回率指标,更全面评估摘要质量
现实中,重大灾难、大型体育赛事等动态事件持续演变。快速获取事件全貌有助于人们及时了解情况并有效应对。然而,关键信息常分散于多篇文档中,涉及复杂的事件知识理解与推理,此前研究对此关注不足。为此,我们提出事件中心型多文档摘要(ECS)任务,旨在基于多篇相关新闻生成简洁全面的事件总结。基于此,我们构建了EventSum数据集,基于百度百科条目并经大规模人工标注,是首个大规模中文多文档摘要数据集,包含5,100个事件、共57,984篇新闻文档,平均每事件有11.4篇输入文档,每事件摘要平均13,471字符。为保障数据质量并避免数据泄露,采用多阶段标注流程对测试集进行人工标注。针对事件信息复杂性,现有评估指标难以全面衡量摘要质量,因此我们设计了事件召回、论据召回、因果召回和时间召回等专用指标及计算方法。我们在EventSum上对先进长上下文大模型进行了全面实验,结果表明:1)现有长上下文大模型在事件中心型摘要任务上仍面临挑战;2)所设计的召回指标对于评估摘要信息完整性至关重要。
原文摘要 · Abstract (English)
In real life, many dynamic events, such as major disasters and large-scale sports events, evolve continuously over time. Obtaining an overview of these events can help people quickly understand the situation and respond more effectively. This is challenging because the key information of the event is often scattered across multiple documents, involving complex event knowledge understanding and reasoning, which is under-explored in previous work. Therefore, we proposed the Event-Centric Multi-Document Summarization (ECS) task, which aims to generate concise and comprehensive summaries of a given event based on multiple related news documents. Based on this, we constructed the EventSum dataset, which was constructed using Baidu Baike entries and underwent extensive human annotation, to facilitate relevant research. It is the first large scale Chinese multi-document summarization dataset, containing 5,100 events and a total of 57,984 news documents, with an average of 11.4 input news documents and 13,471 characters per event. To ensure data quality and mitigate potential data leakage, we adopted a multi-stage annotation approach for manually labeling the test set. Given the complexity of event-related information, existing metrics struggle to comprehensively assess the quality of generated summaries. We designed specific metrics including Event Recall, Argument Recall, Causal Recall, and Temporal Recall along with corresponding calculation methods for evaluation. We conducted comprehensive experiments on EventSum to evaluate the performance of advanced long-context Large Language Models (LLMs) on this task. Our experimental results indicate that: 1) The event-centric multi-document summarization task remains challenging for existing long-context LLMs; 2) The recall metrics we designed are crucial for evaluating the comprehensiveness of the summary information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。