arXiv:2607.11166cs.CL2026-07

构建首个面向主题事件的查询聚焦摘要数据集,解决大规模文档摘要难题。

Query-Focused Event Summarization: A Dataset and Benchmark

论文配图:Query-Focused Event Summarization: A Dataset and Benchmark
图 1 · 摘自论文原文
  • 提出两阶段框架:先自适应检索相关文档,再分层聚类生成摘要。
  • 构建含8个事件、1.67万篇文档、104个查询的QFESum数据集。
  • 适合关注事件追踪与精准信息提取的研究者和应用开发者。

主题语料库是一组语义连贯的文档集合,共同描述一个共享主题事件的不同方面。这类语料库通常包含数百甚至数千篇文档。尽管用户对主题事件的兴趣涉及多个维度,但查询聚焦摘要(QFS)旨在生成符合用户查询的摘要。然而,现有QFS数据集缺乏面向事件的摘要,多数QFS方法难以处理大规模语料库。为应对这些挑战,我们提出查询聚焦事件摘要(QFES)任务,并构建了QFESum数据集,包含8个主题事件、16,684篇文档和104个查询。此外,我们提出一个两阶段QFES框架,包括基于自适应阈值的查询聚焦检索(RAT)和基于分层聚类的查询聚焦摘要(SHC)。在QFESum上的实验结果表明,RAT和SHC持续优于基线模型,证明了其在QFES任务中的有效性。数据集和代码已公开于 https://github.com/sarcasm-hcy02/QFES-QFESum。

原文摘要 · Abstract (English)

A thematic corpus is a collection of semantically coherent documents that collectively describe different aspects of a shared thematic event. Such a corpus typically contains hundreds or even thousands of documents. While users' interests in a thematic event often span multiple dimensions, Query-Focused Summarization (QFS) aims to generate summaries tailored to users' queries. However, existing QFS datasets lack event-oriented summarization, and most QFS methods struggle with large-scale corpora. To address these challenges, we propose the Query-Focused Event Summarization (QFES) task and construct the QFESum dataset, which contains 8 thematic events, 16,684 documents, and 104 queries. Furthermore, we introduce a two-stage QFES framework consisting of Query-Focused Retrieval with Adaptive Thresholding (RAT) and Query-Focused Summarization based on Hierarchical Clustering (SHC). Experimental results on QFESum show that RAT and SHC consistently outperform the baselines, demonstrating their effectiveness for QFES. The dataset and code are publicly available at https://github.com/sarcasm-hcy02/QFES-QFESum.

事件摘要查询聚焦数据集文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。